Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models
机器人操作是视觉到几何映射(f(v)→G):超越语言和视频模型的视觉-几何骨干
机构 * Sun Yat-sen University(中山大学) ; Guangdong Key Laboratory of Big Data Analysis(大数据分析与处理广东省重点实验室) ; X-Era AI Lab(X-Era人工智能实验室) ; AMAP, Alibaba(阿里妈妈实验室) ; Guangdong University of Technology(广东工业大学)
专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.RO;predictive model(abstract)
AI总结 本文提出基于视觉-几何骨干的机器人操作方法,通过直接利用预训练的3D表示生成动作,优于传统语言或视频模型,在精确操作和零样本泛化中表现更优。
Comments Accepted at ACM Multimedia 2026