Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
Agon:具有隐式推理对手评分的竞争性跨模型强化学习
机构 * Independent Researcher(独立研究者)
专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究针对可验证奖励强化学习只评最终答案的问题,提出Agon方法,让两个竞争模型相互评分,通过轮流扮演角色隐式判断推理,在难题上提升了模型表现,且该排序在多种场景和模型家族中可复制。
Comments 15 pages, 7 figures, 8 tables