Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
在做中保留:在线策略数据在缓解遗忘中的作用
专题命中 数学推理 :reasoning(abstract);分类 cs.CL、cs.LG
AI总结 本文系统比较了监督微调(SFT)和强化学习(RL)在语言模型后训练中的遗忘模式,发现RL因使用在线策略数据而更少遗忘,并验证了在线策略数据是缓解遗忘的关键因素。
Journal ref Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026