Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
硬币皆有两面:关于大型语言模型的策略内蒸馏中泛化的双重性质
机构 * University of Science and Technology of China(中国科学技术大学) ; Peking University(北京大学) ; IQuest Research(IQuest研究院) ; MBZUAI(穆罕默德·本·扎耶德人工智能大学) ; Zhejiang University(浙江大学)
AI总结 本研究探究大型语言模型策略内蒸馏(OPD)的泛化特性,发现其迁移教师推理行为而非答案,同源配对泛化性强,异源配对适配性有限,多教师组合存在能力跷跷板效应,为诊断多教师OPD提供了视角。
Comments Under Review