Chimera: Latency- and Performance-Aware Multi-agent Serving for Heterogeneous LLMs
Chimera:面向异构大语言模型的延迟与性能感知多智能体服务
机构 * University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校) ; Microsoft(微软) ; Carnegie Mellon University(卡内基梅隆大学) ; Amazon(亚马逊) ; University of California, Santa Barbara(加州大学圣芭芭拉分校)
AI总结 Chimera通过语义路由和预测调度优化多智能体工作流服务,提升端到端延迟与任务性能,减少1.2-2.4倍延迟并提高8.0-9.5个百分点性能。