HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
HumanVBench: 通过自动合成基准测试探测多模态大语言模型中的以人为本的视频理解
机构 * Sun Yat-Sen University(中山大学) ; Alibaba Group(阿里巴巴集团) ; Peng Cheng Laboratory(鹏城实验室) ; Guangdong Provincial Key Laboratory of Fire Science and Intelligent Emergency Technology(广东省消防科学与智能应急技术重点实验室)
专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 本文提出HumanVBench,一个针对多模态大语言模型(MLLMs)中以人为本的视频理解能力的综合基准测试,通过自动化流程生成高质量视频注释和挑战性问题,揭示30个领先MLLMs在感知细微情绪和对齐语音与视觉线索方面的不足。
Comments Accepted as a conference paper at CVPR 2026