MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
MotionEnhancer: 利用视频扩散模型增强运动感知的视觉-语言模型
机构 * School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) ; Beijing Digital Native Digital City Research Center(北京数字原生数字城研究中心) ; School of Computer Science, Peking University(北京大学计算机学院) ; School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出MotionEnhancer,通过从视频扩散模型中提取运动先验并利用注意力对齐增强视觉-语言模型的运动理解能力,无需额外参数或架构修改,在运动级视频理解基准上取得一致提升。
Comments Accepted by CVPR 2026