Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
从网络视频中隐式几何表示的视觉-语言导航
机构 * Department of Computer Vision, Mohamed Bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed Bin Zayed人工智能大学) ; School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) ; Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区)
专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV
AI总结 本研究提出基于网络视频的隐式几何表示方法,用于视觉-语言导航,通过提升数据利用率和鲁棒性,实现更广泛的应用。
Comments Extension of CVPR 2025 RoomTour3D with implicit geometric representations