CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
CLAP: 从人类视频中学习视觉-语言-动作模型的对比潜在动作预训练
机构 * Tsinghua University(清华大学) ; Astribot ; University of Hong Kong(香港大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 机器人基础模型 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.CV
AI总结 提出CLAP框架,通过对比学习将人类视频与机器人动作词汇对齐,利用伪标签训练VLA模型,实现从人类视频到机器人执行的有效技能迁移。
Comments The code is available at: https://github.com/LinShan-Bin/OpenCLAP