Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map
几何注意力:一种针对Transformer注意力的显式操作语义
机构 * University of Michigan(密歇根大学)
AI总结 几何注意力提出了一种显式操作语义,通过四个独立输入定义注意力层,支持多头、混合核和计划锚等显式领域选择,实现注意力机制的原理性比较与扩展。
Comments 34 pages. Major reconstruction and retitling of the withdrawn previous version. The incorrect non-completability theorem and all dependent claims have been removed. The present version replaces the earlier operator-first development with an analysis of neural-architecture individuation and architecture under composition. Submitted to JMLR