Publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. ECCV
    LISA: Locality-Informed Speculative Decoding for Accelerating Autoregressive Image Generation
    Ying Li, Siyong Jian, Zhaode Wang, Zhiwen Chen, Chengfei Lv, and Huan Wang
    In European Conference on Computer Vision (ECCV), 2026
  2. ECCV
    EVAR: Edge Visual Autoregressive Models via Principled Pruning
    Zefang Wang, Ying Li, Yanyu Li, Mingluo Su, Simin Xu, Guanzhong Tian, and Huan Wang
    In European Conference on Computer Vision (ECCV), 2026
  3. ECCV
    TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
    Kejia Zhang, Keda Tao, Zhiming Luo, Chang Liu, Jiasheng Tang, and Huan Wang
    In European Conference on Computer Vision (ECCV), 2026
  4. ICML
    ARC-Decode: Accelerated Decoding with Risk-Bounded Acceptance
    Ying Li, Zhaode Wang, Zhiwen Chen, Chengfei Lv, and Huan Wang
    In International Conference on Machine Learning (ICML), 2026
  5. ICML
    Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive Generation
    Ying Li, Zefang Wang, Zhaode Wang, Zhiwen Chen, Chengfei Lv, and Huan Wang
    In International Conference on Machine Learning (ICML), 2026
  6. ICML
    Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
    Wenjie Du, Li Jiang, Keda Tao, Xue Liu, and Huan Wang
    In International Conference on Machine Learning (ICML), 2026
  7. ICML
    SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
    Kaiwen Tuo and Huan Wang
    In International Conference on Machine Learning (ICML), 2026
  8. CVPR
    Parallel Jacobi Decoding for Fast Autoregressive Image Generation
    Boya Liao, Ying Li, Siyong Jian, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  9. CVPR
    QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer
    Zhizhen Pan, Hesong Wang, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  10. CVPR
    EarlyTom: Early Token Compression Completes Fast Video Understanding
    Hesong Wang, Xin Jin, Lu Lu, Chenhaowen Li, Jian Chen, Qiang Liu, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  11. CVPR
    OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
    Keda Tao, Kele Shao, Bohan Yu, Weiqiang Wang, Jian Liu, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  12. CVPR
    StreamingTOM: Streaming Token Compression for Efficient Video Understanding
    Xueyi Chen, Keda Tao, Kele Shao, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  13. CVPR
    Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps
    Sicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong, Zikai Song, Jianke Zhu, Huan Wang, and Xinchao Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  14. ICLR
    MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
    Xin Jin, Siyuan Li, Siyong Jian, Kai Yu, and Huan Wang
    In International Conference on Learning Representations (ICLR), 2026
  15. ICLR
    OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
    In International Conference on Learning Representations (ICLR), 2026
  16. ICLR
    RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
    Sicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong, Jianke Zhu, and Huan Wang
    In International Conference on Learning Representations (ICLR), 2026
  17. ICLR
    Autoregressive Image Generation with Randomized Parallel Decoding
    Haopeng Li, Jinyue Yang, Guoqi Li, and Huan Wang
    In International Conference on Learning Representations (ICLR), 2026
  18. CPAL Oral
    ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning
    Mingluo Su and Huan Wang
    In Conference on Parsimony and Learning (CPAL), 2026
  19. CPAL
    ResSVD: Residual Compensated SVD for Large Language Model Compression
    Haolei Bai, Siyong Jian, Tuo Liang, Yu Yin, and Huan Wang
    In Conference on Parsimony and Learning (CPAL), 2026
  20. TMLR
    When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios
    Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, and Huan Wang
    Transactions on Machine Learning Research (TMLR), 2026
  21. TMLR
    Is Oracle Pruning the True Oracle?
    Sicheng Feng, Keda Tao, and Huan Wang
    Transactions on Machine Learning Research (TMLR), 2026
  22. arXiv
    LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
    Keda Tao, Yuhua Zheng, Jia Xu, Wenjie Du, Kele Shao, Hesong Wang, Xueyi Chen, Xin Jin, Junhan Zhu, Bohan Yu, Weiqiang Wang, Jian Liu, Can Qin, Yulun Zhang, Ming-Hsuan Yang, and Huan Wang
    arXiv preprint arXiv:2603.19217, 2026
  23. arXiv
    MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
    Xingze Zou, Jing Wang, Yuhua Zheng, Xueyi Chen, Haolei Bai, Lingcheng Kong, Syed A.R. Abu-Bakar, Zhaode Wang, Chengfei Lv, Haoji Hu, and Huan Wang
    arXiv preprint arXiv:2603.11935, 2026
  24. arXiv
    DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
    Haolei Bai, Lingcheng Kong, Xueyi Chen, Jianmian Wang, Zhiqiang Tao, and Huan Wang
    arXiv preprint arXiv:2602.11715, 2026

2025

  1. NeurIPS
    FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware Guidance
    Ying Li, Chengfei Lv, and Huan Wang
    In Advances in Neural Information Processing Systems (NeurIPS), 2025
  2. NeurIPS
    HoliTom: Holistic Token Merging for Fast Video Large Language Models
    Kele Shao, Keda Tao, Can Qin, Haoxuan You, Yang Sui, and Huan Wang
    In Advances in Neural Information Processing Systems (NeurIPS), 2025
  3. NeurIPS
    Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
    Kejia Zhang, Keda Tao, Jiasheng Tang, and Huan Wang
    In Advances in Neural Information Processing Systems (NeurIPS), 2025
  4. CVPR
    DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
    Keda Tao, Can Qin, Haoxuan You, Yang Sui, and Huan Wang
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  5. TCSVT
    Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
    Xianzu Wu, Zhenxin Ai, Harry Yang, Ser-Nam Lim, Jun Liu, and Huan Wang
    IEEE Transactions on Circuits and Systems for Video Technology, 2025
  6. arXiv
    Active Perception Agent for Omnimodal Audio-Video Understanding
    Keda Tao, Wenjie Du, Bohan Yu, Weiqiang Wang, Jian Liu, and Huan Wang
    arXiv preprint arXiv:2512.23646, 2025
  7. arXiv
    ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
    Lingcheng Kong, Jiateng Wei, Hanzhang Shen, and Huan Wang
    arXiv preprint arXiv:2510.07356, 2025
  8. arXiv
    Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
    Keda Tao, Haoxuan You, Yang Sui, Can Qin, and Huan Wang
    arXiv preprint arXiv:2503.16257, 2025