Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
基于层级优势加权的在线RL微调VLA策略从稀疏回合结果
机构 * ACE Robotics ; Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 VLA模型 :VLA(title_cn,summary_cn);分类 cs.RO、cs.LG
AI总结 提出层级优势加权行为克隆(HABC),通过分离生存性和效率目标并自适应平衡,解决稀疏二元结果下VLA策略在线微调中的信用分配问题,在三个双臂接触任务上将成功率从12-44%提升至38-92%。
Comments Website: https://acerobotics-vla.github.io/HABC-Website