VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
VLAFlow:通过协同训练和未来潜在对齐的视觉-语言-动作模型统一训练框架
Guoyang Xia, Fengfa Li, Hongjin Ji, Lei Ren, Fangxiang Feng, Kun Zhan, Yan Xie
机构
*
Li Auto Inc.(理想汽车)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
Comments37 pages, 2 figures. Any comments are welcomed. v2: Significant revisions on arguments, conclusion unchanged. An action model for covariant observer added. v3: Version submitted to JHEP. Clarifications and references added