Incentivizing Vision Language Models to Search for Long Video Question Answering
激励视觉语言模型进行长视频问答搜索
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 视频问答 :long video(title,abstract);video understanding(abstract);分类 cs.CV
AI总结 研究将长视频问答从被动单流程感知任务转变为多轮检索过程,核心方法是自然语言驱动搜索及强化学习后训练,贡献是提升长视频理解基准测试分数。
Comments To appear at the European Conference on Computer Vision (ECCV 2026)