VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
VLAFlow:通过协同训练和未来潜在对齐的视觉-语言-动作模型统一训练框架
机构 * Li Auto Inc.(理想汽车) ; School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
AI总结 提出VLAFlow统一框架,比较四种训练范式,发现语言监督保持视觉-语言泛化,未来潜在对齐改进状态转换建模,两者结合实现最佳迁移性能。