Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
用奖励模型增强大语言模型推理能力:一项分析性综述
机构 * Department of Data Science and Hong Kong Institute of AI for Science, City University of Hong Kong(数据科学系和香港人工智能科学研究所,香港城市大学) ; Li Auto Inc., China(中国利汽车公司) ; Department of Statistics, University of Oxford(统计系,牛津大学)
AI总结 该综述系统介绍奖励模型(RMs)的概念、训练与评估方法,梳理其在大语言模型(LLMs)推理中的三类核心应用,并探讨RMs相关开放问题,为其有效部署提供见解。
Comments Accepted for publication in Artificial Intelligence Review