RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
RoboRefer: 向视觉语言模型在机器人中的空间指称推理迈进
机构 * School of Software, Beihang University(北京航空航天大学软件学院) ; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院)
专题命中 视觉推理 :vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
AI总结 RoboRefer通过整合深度编码器和强化微调方法,实现了视觉语言模型在机器人中的空间指称推理,提升了复杂场景下的交互能力。
Comments Accepted by NeurIPS 2025. Project page: https://zhoues.github.io/RoboRefer/