Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
机构 * Tsinghua University(清华大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 预训练与数据 :pretraining(title,abstract);分类 cs.AI、cs.LG
Comments Project page: https://microsoft.github.io/VITRA/