arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Learning Representations · 会议 · Machine Learning

2026-07-31 至 2026-07-31 共收录 3
2607.27987 2026-07-31 cs.LG 新提交

It's All Just Vectorization: einx, a Universal Notation for Tensor Operations

这全都是向量化:einx,一种通用的张量运算表示法

Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens

AI总结 针对主流张量框架表示法难读写、易出错的问题,引入通用张量运算表示法einx,简化API、统一规则,提供可与现有框架无缝集成的Python实现。

Comments Published at ICLR 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28259 2026-07-31 cs.LG math.AT 新提交

TopoFormer: Topology Meets Attention for Graph Learning

TopoFormer:拓扑与注意力融合的图学习方法

Md Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora, Baris Coskunuzer

AI总结 该研究提出TopoFormer框架,通过核心模块Topo-Scan将图拓扑结构编码为注意力友好序列,结合Transformer实现图表示学习,在相关基准上达到SOTA性能,开辟了拓扑与注意力融合的图学习新方向。

Comments 26 pages, 5 figures

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05554 2026-07-31 cs.LG cs.AI cs.DM math.CA 版本更新

Critical attention scaling in long-context transformers

长上下文Transformer中的关键注意力缩放

Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet

机构 * Department of Mathematics, Massachusetts Institute of Technology(数学系,麻省理工学院) Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(电气工程与计算机科学系,麻省理工学院)

AI总结 该研究针对长上下文Transformer的注意力秩崩溃问题,通过分析简化模型确定关键缩放因子$\beta_n \backsim \text{log} n$,为YaRN和Qwen的注意力缩放提供理论依据,阐明对数缩放可维持长上下文下的稀疏内容自适应注意力。

Comments 31 pages, 2 figures

Journal ref Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏