Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
用评分奖励模型治愈大语言模型数学推理中的奇迹步骤
机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen, China(数据科学学院,香港中文大学(深圳)) ; UC Berkeley(加州大学伯克利分校) ; Zhejiang University(浙江大学) ; Johns Hopkins University(约翰霍普金斯大学) ; Renmin University of China(中国人民大学) ; Xiaohongshu Inc.(小红书公司)
AI总结 本文通过评分奖励模型解决大语言模型数学推理中的奇迹步骤问题,通过系统分析和人类验证建立失败模式分类,提升推理准确性和可靠性。
Comments Accepted by ACL 2026 Main, 22 pages, 10 figures, 7 Tables