LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
LEO-VL:面向可扩展3D视觉-语言学习的高效场景表示
机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) ; School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) ; Department of Automation, Tsinghua University(清华大学自动化系)
专题命中 其他3D视觉 :3D vision(title,abstract);分类 cs.CV
AI总结 本文提出LEO-VL,一种基于高效场景表示CFG的3D视觉-语言模型,通过改进训练数据和引入SceneDPO提升鲁棒性,在多个3D-VL基准测试中取得最佳性能。
Comments Project page: https://leo-vl.github.io