Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
解耦推理与置信度:在可验证奖励的强化学习中恢复校准
机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院软件研究所信息处理实验室) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) ; Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院) ; National Computer Network Emergency Response Technical Team/Coordination Center of China, Beijing, China(中国国家计算机网络应急技术配合中心)
AI总结 针对RLVR中模型校准退化问题,提出DCPO框架通过解耦推理与校准目标,在保持准确率的同时显著改善校准性能并缓解过度自信。
Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)