AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
AgentOPSD:面向智能体强化学习的递归自蒸馏方法
机构 * Tsinghua University(清华大学) ; Zhejiang University(浙江大学) ; Meituan(美团)
AI总结 提出 AgentOPSD 这一无 critic 的递归自蒸馏方法,用于智能体强化学习的轮次级 credit 分配,在 ALFWorld 等数据集上优于 GRPO 等基线,在 Qwen2.5-7B 模型上实现 ALFWorld 89.1% 的成功率。
Comments Code: https://github.com/ZethWang/AgentOPSD