What should post-training optimize? A test-time scaling law perspective
在测试时间视角下,后训练应优化什么?
机构 * Department of Statistical Sciences, University of Toronto(多伦多大学统计科学系) ; Department of AI and Data Science, The University of Hong Kong(香港大学人工智能与数据科学系)
AI总结 本文研究了测试时间预算不匹配场景下,基于奖励尾部统计的后训练优化方法,提出TEA和Prefix-TEA估计器提升最佳N性能。