Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
通过真实世界室内游览视频的多模态事件知识增强视觉语言导航
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Tsinghua University(清华大学) ; Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(马尔代夫 bin Zayed 大学人工智能学院) ; Hangzhou City University(杭州城市学院)
AI总结 本文提出基于多模态事件知识的视觉语言导航增强方法,通过构建大规模时空知识图谱并结合层次检索机制,提升长视界推理和粗粒度指令处理能力。