Linear Dynamics in the RLVR Training of Large Language Models
在大语言模型RLVR训练中的线性动力学
机构 * Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) ; Hong Kong Institute of AI for Science, City University of Hong Kong(香港城市大学人工智能科学研究院) ; Li Auto Inc. ; Beihang University(北航大学)
AI总结 本文研究了强化学习可验证奖励(RLVR)在大语言模型训练中的内部动态,发现RLVR在多种模型和训练配置下均进入线性区域,通过实验和理论分析证明这种线性特性源于训练信号的高方差和噪声,且具有预测性和实用性。
Comments Major revision: substantially reorganized the manuscript and added a theoretical explanation section. The replacement is intended for the same arXiv paper; the core topic and contribution remain the same