How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
大语言模型和视觉语言模型如何在没有视觉信息的情况下理解视角旋转?一项可解释性研究
机构 * School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) ; Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(教育部计算功率网络与信息安全部重点实验室,山东计算机科学中心(济南国家超级计算中心),齐鲁工业大学(山东省科学院))
AI总结 本文研究了在无视觉信息情况下,大语言模型和视觉语言模型理解视角旋转的能力,发现两者表现不佳,而人类可达到100%准确率,揭示了模型与空间智能需求之间的差距。
Comments Published as a main-conference paper at The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)