When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
何时学习停止有帮助?推理模型中早期退出的成本感知研究
机构 * University of Maine at Presque Isle(缅因大学普雷斯克岛分校) ; Stanford University(斯坦福大学) ; Independent Researcher(独立研究员)
专题命中 数学推理 :reasoning(title,abstract);test-time compute(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究推理语言模型中学习停止规则相对于简单置信度或收敛阈值的优势,提出LearnStop方法,发现其效果取决于任务类型,在自由形式数学任务中表现优于标量退出。
Comments 23 pages, 5 figures