XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
机构 * Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国) ; School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国) ; State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国) ; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract)
AI总结 XR-1通过学习统一的视觉-运动表示,解决视觉-语言-动作模型在低级动作生成和跨数据源领域差距的挑战,提出三阶段训练方法并验证了其在多种机器人和任务上的优越性能。
Comments Accepted to ICML2026 as Oral