RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
RLinf-VLA: 一种用于视觉-语言-动作模型强化学习的统一高效框架
机构 * Tsinghua University(清华大学) ; Zhongguancun Academy(中关村学院) ; Infinigence AI ; Peking University(北京大学) ; UC Berkeley(加州大学伯克利分校) ; Harbin Institute of Technology(哈尔滨工程学院) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
专题命中 VLA模型 :VLA(title,title_cn);vision-language-action(title,abstract);action model(title);分类 cs.RO
AI总结 RLinf-VLA是一种统一高效的框架,用于提升视觉-语言-动作模型在具身环境中的强化学习性能,通过统一接口和高效资源分配实现统一和可扩展的训练。
Comments Accepted to RSS 2026. This is the technical report of the RLinf Team, focusing on the algorithm side. For the system-level design, please refer to arXiv:2509.15965. The open-sourced code link: https://github.com/RLinf/RLinf