The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
视觉-语言模型中空间变量绑定的双重机制
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) ; Northeastern University(东北大学) ; Sony Playstation(索尼PlayStation)
专题命中 视觉问答 :vision-language model(title,abstract);VLM(abstract_cn);visual question answering(abstract);分类 cs.CV、cs.LG
AI总结 本文揭示视觉-语言模型通过语言骨干中的内容无关空间关系编码和视觉编码器中的全局布局表示两种机制实现空间变量绑定,其中视觉编码器起主导作用。
Comments 37 pages, 53 figures