Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Uni-OPD:基于双视角方案的统一策略内蒸馏框架
机构 * Zhejiang University(浙江大学) ; Shenzhen Loop Area Institute(深圳河套学院) ; LLM Department, Tencent(腾讯大模型部门)
AI总结 针对策略内蒸馏存在的信息状态探索不足、教师监督不可靠问题,提出跨大语言与多模态大模型的统一框架Uni-OPD,经多域实验验证其通用性与有效性。