Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
看见不等于共享:一些视觉-语言模型在非对称对话中高估共同基础
机构 * Utrecht University(乌得勒支大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract_cn);grounding(abstract);分类 cs.AI
AI总结 研究视觉-语言模型在对话中是否混淆潜在共享与已共享信息,发现模型过度依赖地图内容预测对齐,忽视对话历史中的基础建立过程。
Comments 17 pages, 9 figures, 8 tables; accepted to SIGDIAL 2026