SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
SPIRAL:通过多智能体多轮强化学习进行零和游戏的自我对战以促进推理
机构 * National University of Singapore(新加坡国立大学) ; Northeastern University(东北大学) ; Sea AI Lab(Sea AI 实验室) ; Centre for Frontier AI Research (CFAR), A*STAR(前沿人工智能研究中心(CFAR),A*STAR) ; Plastic Labs ; University of Washington(华盛顿大学)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 SPIRAL通过多智能体多轮强化学习在零和游戏中促进推理,展示了模型在多个基准测试中的显著性能提升。
Comments Accepted at ICLR 2026. Code: https://github.com/spiral-rl/spiral