GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
GA-VLN: 用于高效视觉-语言导航的几何感知鸟瞰图表示
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Robbyant ; School of Computing, National University of Singapore(新加坡国立大学计算机学院) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 BEV与占用 :BEV(title,summary_cn);分类 cs.CV、cs.AI
AI总结 本文提出GA-VLN框架,通过引入几何感知的鸟瞰图表示(GA-BEV),整合显式和隐式几何信息,提升视觉-语言导航的效率和性能,实验表明其在仅使用导航数据的情况下取得了最先进的结果。