Throughput-Optimal Scheduling Algorithms for LLM Inference and AI Agents
面向LLM推理和AI代理的吞吐量最优调度算法
机构 * School of Operations Research and Information Engineering, Cornell University(Cornell大学运筹学与信息工程学院) ; Operations Management, Booth School of Business, University of Chicago(芝加哥大学博斯商学院运营管理系) ; Decision, Risk and Operations, Columbia Business School, Columbia University(哥伦比亚大学哥伦比亚商学院决策、风险与运营系)
专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 本文从排队论角度研究了LLM推理系统的吞吐量优化问题,证明了工作保持调度算法在DAG和Fork-Join路由拓扑中能实现最大吞吐量,并揭示了批量处理网络中K-FCFS调度的流极限框架,评估了Orca和Sarathi-Serve的吞吐量最优性,同时指出批量大小限制和循环路由拓扑对吞吐量的影响。