S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
S2-MoE:在边缘设备上实现混合专家模型的高效自推测解码
机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) ; School of Integrated Circuits, Peking University(北京大学集成电路学院) ; School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
AI总结 S2-MoE是面向边缘设备MoE推理的高效自推测解码框架,通过路由感知自适应扩展、复用感知门控等设计,使MoE推理在边缘设备上最高获5.3倍加速,平均约2.0倍。
Comments 13 pages, 10 figures