Rethinking LLM Ensembling from the Perspective of Mixture Models
从混合模型的角度重新思考大语言模型集成
机构 * Key Laboratory of New Generation Artificial Intelligence Technology(新一代人工智能技术关键实验室) ; Its Interdisciplinary Applications (Southeast University), Ministry of Education(交叉应用(东南大学),教育部) ; Southeast University(东南大学) ; Centre for Frontier AI Research (CFAR), Agency for Science, Technology(前沿人工智能研究(CFAR),科技研究局) ; Research (A STAR), Singapore(研究(A STAR),新加坡) ; Institute of High Performance Computing (IHPC), Agency for Science, Technology(高性能计算(IHPC),科技研究局)
AI总结 本文提出混合模型式集成(ME),通过将集成重新解释为混合模型,随机选择单个模型生成下一个token,避免显式计算完整集成分布,实现1.78x-2.68x加速,并揭示了集成与token级路由方法的联系。
Comments ICML 2026 Spotlight