Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
教模型自学:在可学习性边缘推理
机构 * MIT(麻省理工学院) ; Meta FAIR ; New York University(纽约大学)
AI总结 提出SOAR非对称自对弈框架,通过元强化学习生成自动课程,解决低初始成功率数据集上的推理模型扩展问题,发现基于学生进步的奖励优于内在可学习性奖励,且问题结构比答案正确性更重要。
Comments ICML 2026. Blog post: https://ssundaram21.github.io/soar/