Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
通过训练无关的置信度感知校准提高扩散式大语言模型的吞吐量
机构 * Rice University(里士大学) ; Intel(英特尔) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校)
AI总结 本文提出CadLLM,一种无需训练的加速扩散式大语言模型(dLLMs)推理吞吐量的方法。通过动态调整生成块大小、步长和阈值,结合置信度控制,减少softmax开销,提升吞吐量。
Comments 12 pages, 3 figures. Accepted to Findings of ACL 2026