Critical attention scaling in long-context transformers
长上下文Transformer中的关键注意力缩放
机构 * Department of Mathematics, Massachusetts Institute of Technology(数学系,麻省理工学院) ; Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(电气工程与计算机科学系,麻省理工学院)
AI总结 该研究针对长上下文Transformer的注意力秩崩溃问题,通过分析简化模型确定关键缩放因子$\beta_n \backsim \text{log} n$,为YaRN和Qwen的注意力缩放提供理论依据,阐明对数缩放可维持长上下文下的稀疏内容自适应注意力。
Comments 31 pages, 2 figures
Journal ref Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026), 2026