Watch Before You Answer: Learning from Visually Grounded Post-Training
在回答前观看:从视觉引导的后训练中学习
机构 * University of British Columbia(不列颠哥伦比亚大学) ; Vector Institute(向量研究所) ; Etude AI ; Kolors Team, Kuaishou Technology(快手科技Kolors团队) ; University of Toronto(多伦多大学) ; University of Waterloo(滑铁卢大学) ; University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
AI总结 本文发现现有视频理解基准中40-60%的问题可通过文本线索回答,提出VidGround方法通过仅使用视觉引导问题提升VLM性能,实验显示其在后训练中效果优于复杂方法。