Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
通过自我场景增强在多模态大语言模型中强化自我中心空间感知
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 研究如何强化多模态大语言模型的自我中心空间感知,提出自我场景增强框架ESA,利用自我元素图作为中间表示,通过视觉基础模型增强空间感知,在EgoTextVQA基准上取得显著性能提升。
Comments 14 pages, 8 figures. Chi Kit Wong and Ye Pan contributed equally. Code: this https URL (https://github.com/Chikit-WONG/spatialGraph)