On-Policy Self-Distillation without Any Supervision
无任何监督的在线策略自蒸馏
机构 * UC San Diego(加州大学圣迭戈分校) ; Georgia Institute of Technology(佐治亚理工学院) ; University of Maryland, College Park(马里兰大学帕克分校) ; ByteDance(字节跳动)
AI总结 本研究提出无监督在线策略自蒸馏(U-OPSD),仅用模型自身生成结果实现在线策略自蒸馏,在多数学基准上优于基础模型,部分场景超越OPSD、GRPO等监督方法。
Comments Project page at https://williamium3000.github.io/u-opsd/ and code at https://github.com/williamium3000/u-opsd