What We are Missing in Multimodal LLM Evaluation?
我们在多模态大语言模型评估中缺失了什么?
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI
AI总结 本文分析现有多模态大语言模型评估基准的不足,识别出时空连贯性、物理世界理解、多模态一致性和选择性注意等关键缺失,强调填补这些空白对衡量多模态智能真实进展的重要性。