Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought
CoT-Space: 一种通过强化学习实现内部慢思考的理论框架
机构 * Zeyu Gan, Yi Hao, Yong Liu(GAN 赵毅、LIU 刘永)
专题命中 复杂问题求解 :CoT(title_cn,summary_cn);reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出CoT-Space理论框架,通过强化学习将推理过程从离散的token预测任务转化为连续的推理层面语义空间中的优化过程,揭示了测试时扩展中最优CoT长度的收敛是欠拟合与过拟合基本权衡的自然结果。
Comments Preprint Edition