Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
注意力-前馈网络解耦大语言模型服务的分析资源配置
机构 * Dept. of Industrial Engineering and Decision Analytics HKUST(工业工程与决策分析系香港科技大学) ; Dept. of Computer Science and Technology Tsinghua University(计算机科学与技术系清华大学) ; IIIS Tsinghua University(清华大学信息学院) ; Huawei Hong Kong Research Center(华为香港研发中心) ; School of Mathematical Sciences Peking University(北京大学数学科学学院)
AI总结 本文提出在随机负载下,针对注意力-前馈网络解耦架构的分析资源配置框架,通过考虑工作负载统计量θ,确定最优注意力与前馈网络比例,减少阻塞和设备空闲时间。
Comments Submitted to Neurips 2026