Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
为何扩散语言模型在真正并行(非自回归)解码中表现不佳?
机构 * The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) ; ELLIS Institute Tübingen, Tübingen, Germany(图宾根ELLIS研究所) ; Max Planck Institute for Intelligent Systems, Tübingen, Germany(智能系统马克斯·普朗克研究所) ; Tübingen AI Center, Tübingen, Germany(图宾根人工智能中心) ; University of Surrey, Guildford, United Kingdom(萨里大学) ; The University of North Carolina at Chapel Hill, Chapel Hill, NC, USA(北卡罗来纳大学教堂山分校)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);math reasoning(abstract)
AI总结 本文提出NAP方法,通过数据驱动策略改进扩散语言模型的非自回归并行解码性能。