ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
ViSTR-Bench:多模态大语言模型能否从动态场景中的连续视觉线索进行推理?
机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) ; Zhongguancun Academy(中关村科学城) ; School of Engineering, Westlake University(西湖大学工学院) ; School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
AI总结 研究聚焦多模态大语言模型在动态场景中从连续视觉线索推理的能力,提出ViSTR-Bench评估套件,基于相关原则建立四维评估,含15个子任务和1340个问答对,经评估发现当前模型在复杂时空推理上有瓶颈,远不及人类。
Comments 37 pages, 37 figures