Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
大语言模型推理服务系统的比较分析:vLLM和HuggingFace TGI的性能研究
机构 * Baltimore, MD, USA(美国马里兰州巴尔的摩)
专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,comments);分类 cs.LG
AI总结 本文比较了vLLM和HuggingFace TGI在大语言模型推理服务中的性能,发现vLLM在高并发场景下吞吐量更高,而TGI在交互式应用中延迟更低。
Comments 10 pages, benchmarking study of LLM inference systems