Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
分层验证投机波束以加速大语言模型推理
机构 * Praxis Business School(普拉克斯商学院) ; Sabre Industries(Sabre工业公司)
专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 提出分层验证树(HVT)框架,通过优先验证高似然草稿并早期剪枝次优候选,以分层方式重构投机波束解码,从而在不重训练或修改架构下显著降低推理时间和能耗。
Comments This paper was accepted for oral presentation and publication in the 3rd International Conference on Data Science and Network Engineering (ICDSNE 2025), organized at NIT, Agartala, India, from July 25 to 26, 2025. The paper is 12 pages long, and it contains 3 tables and 4 figures. This is NOT the final paper, which will be published in the Springer-published proceedings