From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
从推理到代理:大型语言模型强化学习中的信用分配
机构 * Independent Researcher(独立研究员)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 本文探讨了大型语言模型强化学习中信用分配问题,总结了47种方法,并提出了可重用的资源,指出从推理到代理的转变使信用分配更加复杂,推动了新的方法发展。
Comments 50 pages, 4 figures. v3: expanded to a unified 69-paper corpus through July 31, 2026; adds restored-state identification results, blind coding of a 42-paper full-text subset, a source-located reporting audit, the CA-ID Card, and replay-fidelity analysis