PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
PanoGrounder: 利用全景场景表示桥接2D和3D,实现基于VLM的3D视觉定位
机构 * Seoul National University(首尔大学) ; Robotics Lab, Hyundai Motor Company(现代汽车公司机器人实验室) ; Pohang University of Science and Technology (POSTECH)(浦项科技大学)
专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);vision-language model(abstract);分类 cs.CV
AI总结 提出PanoGrounder框架,通过多模态全景表示与预训练2D VLM结合,实现强泛化能力的3D视觉定位,在ScanRefer和Nr3D上取得最优结果。
Comments ECCV 2026