Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
重新审视在线策略蒸馏:经验性失败模式及简单修复
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中科院自动化所) ; School of Artificial Intelligence, UCAS(中国科学技术大学人工智能学院) ; Fudan University(复旦大学) ; Independent Researcher(独立研究者)
专题命中 效率与部署 :LLM(abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文研究了在线策略蒸馏的失败模式,提出改进方法提升训练稳定性,实验显示改进方法在多个任务中提升性能19.8%。