Accuracy-Delay Trade-Off in LLM Offloading via Token-Level Uncertainty
在LLM卸载中通过令牌级不确定性实现精度-延迟权衡
机构 * Dept. of Electrical and Computer Engineering, Seoul National University, Seoul, Korea(电子与计算机工程系,首尔国立大学,首尔,韩国) ; Institute of New Media and Communications, Seoul National University, Seoul, South Korea(新媒体与通讯研究所,首尔国立大学,首尔,韩国)
专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 本文提出了一种基于令牌级不确定性的卸载框架,通过贪心算法在保持准确性的同时减少延迟,实现了LLM在移动边缘计算中的精度-延迟权衡。
Comments This paper has been accepted at 2025 IEEE Globecom Workshop: WS02-GAIMC: Mutual Facilitation of Generative Artificial Intelligence and Mobile Communications