CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment
CLAR: 通过融合掩码重建与多层级对比对齐学习用于机器人操作的3D表示
机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所SKL-MAIS) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Carnegie Mellon University(卡内基梅隆大学) ; Galbot ; CFCS, School of Computer Science, Peking University(北京大学计算机科学与技术学院CFCS)
AI总结 提出CLAR框架,融合掩码自编码与全局跨模态对比学习,并引入基于可变形注意力的局部自适应对齐机制,解决3D预训练中空间几何与语义细节的权衡问题,在视觉运动策略学习中达到最优性能。