VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
VANE:通过未来视觉表示预测实现视觉-语言-动作模型的可靠测试时训练
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Beijing University of Posts and Telecommunications(北京邮电大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Li Auto Inc.(理想汽车)
AI总结 本文提出VANE框架,通过条件提示自适应、未来视觉后果学习及候选更新隔离评估的方法,提升VLA策略在闭环操纵中的可靠性,在SimplerEnv WidowX上将平均成功率提高3.2个百分点。