AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
AAPO: 通过优势边际增强大语言模型的推理能力
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Frontier Research Department, Baidu Inc.(百度公司前沿研究部)
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL、cs.LG
AI总结 AAPO通过引入基于边际的估计方案优化交叉熵损失,有效解决传统分组相对优势估计方法的训练效率问题,在数学推理基准上表现出色。
Comments Accepted to ACL2026 Main Conference