arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Transactions on Machine Learning Research · 期刊 · Machine Learning

2026-06-25 至 2026-06-25 共收录 2
2606.24975 2026-06-25 cs.LG cs.AI cs.CL 新提交

Why Do Accumulated Transformations Extrapolate?

为什么累积变换能够外推?

Mahesh Godavarti

机构 * A Carrot, Inc.(A Carrot公司)

AI总结 本文研究累积正交变换(如Householder反射或SO(2)旋转)在注意力机制中产生长度外推能力的原理,证明其通过有限步后去相干性抑制远距离token,并指出其最终会退化,而旋转值可扩展有效范围。

Comments 33 pages, submitted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24970 2026-06-25 cs.LG 新提交

Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

不要破坏我的LLM:剪枝注意力层对解释忠实性和置信度校准的影响

Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo, Christina Lioma

机构 * University of Copenhagen(哥本哈根大学) LUT University(拉赫蒂理工大学)

AI总结 研究剪枝LLM注意力层对解释忠实性和置信度校准的影响,发现尽管准确率保持,但忠实性和校准度常下降,表明模型置信度、可解释性与准确性之间存在错位。

Comments Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏