MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
MLA:一种多感官语言-动作模型,用于机器人操作中的多模态理解和预测
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) ; Beijing Innovation Center of Humanoid Robotics (X-Humanoid)(北京类人机器人创新中心(X-Humanoid)) ; The Chinese University of Hong Kong (CUHK)(香港中文大学(CUHK))
专题命中 机器人操作 :robotic(title,abstract);manipulation(title,abstract);world model(abstract);分类 cs.RO
AI总结 本文提出MLA模型,通过多感官融合与未来目标预测,提升机器人在复杂接触任务中的操控能力,实验显示其在2D和3D任务中性能优于现有方法。
Comments Project page: https://robotic-mla.github.io/