机构
*
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所模式识别实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Kling Team, Kuaishou Technology(快手科技 Kling 团队)
;
Peking University(北京大学)
;
Nanjing University(南京大学)
机构
*
School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)
;
Research Center for SCIR, Harbin Institute of Technology(SCIR研究中心,哈尔滨工业大学)
;
NExT Research Center, National University of Singapore(NExT研究中心,新加坡国立大学)
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Minghao Qin, Xiangrui Liu, Zhengyang Liang, Yan Shu, Huaying Yuan, Juenjie Zhou, Shitao Xiao, Bo Zhao, Zheng Liu
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Trento(特伦特大学)
;
Renmin University of China(中国人民大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Hong Kong Polytechnic University(香港理工大学)
机构
*
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Guizhou Provincial Laboratory of Big Data, College of Computer Science and Technology, Guizhou University(贵州大数据省实验室,贵州大学计算机科学与技术学院)
Comments14 pages main manuscript with 3 figures; 6 pages supplementary material with 3 figures. To be presented at International Conference on Information Processing in Computer-Assisted Interventions (IPCAI 2025). To be published in International Journal of Computer Assisted Radiology and Surgery (IJCARS)