Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models
超越答案的思考:评估大型推理模型中的有害过度思考
机构 * University of Trento(特伦托大学) ; Toyota Motor Europe(丰田欧洲公司) ; Fondazione Bruno Kessler(布鲁诺·凯塞林基金会)
专题命中 测试时计算 :reasoning(title,abstract);test-time compute(abstract);分类 cs.AI
AI总结 本文提出前缀级轨迹评估协议,通过定义推理充分性来区分冗余但无害的冗长过度思考和导致正确轨迹偏离的有害过度思考,发现当前模型不仅受限于推理能力,还受限于无法在适当时机停止。