Weak-to-Strong Generalization via Direct On-Policy Distillation
通过直接在线策略蒸馏实现弱到强泛化
机构 * SIA-Lab of Tsinghua AIR and ByteDance Seed(清华-字节跳动联合研究中心SIA实验室) ; Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) ; Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; Peking University(北京大学)
AI总结 针对可验证奖励RL在大模型上成本高的问题,提出Direct-OPD方法迁移小模型RL的策略偏移,无需在大模型上跑RL即可实现性能提升。
Comments Project Page: https://bytedtsinghua-sia.github.io/Direct-OPD/