FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning
FastMMoE:通过动态专家激活和路由感知的标记剪枝加速多模态大语言模型
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Li Auto(利亚自动化)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
AI总结 本文提出FastMMoE,一种无需训练的加速框架,通过动态专家激活和路由感知标记剪枝,显著降低计算量并保持性能,优于现有基线方法。