ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
ImagineNav++: 通过场景想象促使视觉语言模型作为具身导航器
机构 * School of Automation, Southeast University(东南大学自动化学院)
专题命中 视觉空间推理 :reasoning(abstract);planning(abstract)
AI总结 本文提出ImagineNav++,通过场景想象将视觉语言模型用于无地图导航,利用想象模块生成高探索潜力的视点,并通过选择性聚焦记忆机制实现空间一致性,实验表明其在无地图导航中表现优异。
Comments 17 pages, 10 figures. arXiv admin note: text overlap with arXiv:2410.09874