ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
ResRL: 通过负样本投影残差强化学习提升大语言模型推理能力
机构 * MAIS\&NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(MAIS与NLPR,自动化研究所,中国科学院,北京,中国) ; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China(先进交叉学科学院,中国科学院大学,北京,中国)
AI总结 本文提出ResRL,通过解耦正负响应的语义分布,提升LLM推理能力同时保持生成多样性,在十二个基准测试中超越强基线,尤其在数学推理上表现更优。
Comments Accepted to ICML 2026. Preprint version. https://github.com/1229095296/ResRL.git