In-context superposition: human-like working memory interference in large language models
类人工作记忆干扰在大语言模型中
Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
机构
*
School of Psychological and Brain Sciences, Georgia Tech(佐治亚理工学院心理与脑科学学院)
;
Department of Psychology, New York University(纽约大学心理学系)
;
Department of Cognitive Science, Indiana University Bloomington(印第安纳大学布卢明顿分校认知科学系)
;
Honda Research Institute(本田研究所)
;
Center of Excellence for Computational Cognition, Georgia Tech(佐治亚理工学院计算认知卓越中心)
;
Departments of Neuroscience and Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校神经科学和心理学系)
专题命中
长上下文与记忆
:large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
ERSkill:面向技能引导的自适应记忆检索的演化框架
Haolong Chen, Liang Zhang, Zhuo Li, Lei Xue, Guanrxu Zhu
机构
*
Shenzhen International Center for Industrial and Applied Mathematics(深圳国际工业与应用数学中心)
;
Shenzhen Research Institute of Big Data(深圳大数据研究院)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
;
Shenzhen Loop Area Institute(深圳河套学院)
专题命中
长上下文与记忆
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models
类人短暂记忆提升语言学习但损害变压器语言模型的阅读时间预测
Abishek Thamma, Micha Heilbron
机构
*
University of Amsterdam, Amsterdam Brain and Cognition(阿姆斯特丹大学,阿姆斯特丹脑与认知中心)
;
Vrije Universiteit Amsterdam, Department of Informatics(阿姆斯特丹自由大学,信息学院)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学语言学研究所)
Commentsv2: Revised after peer review. Accepted for publication in Transactions of the Association for Computational Linguistics v3: Added link to code repository. Code: https://github.com/drhanjones/fmt-llm
Journal refTransactions of the Association for Computational Linguistics 14 (2026) 877-892
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
预测随机低维重参数化何时能训练神经网络
Andrew Cheng, Ali Eslamian, Jie Cheng, Mehdi Zargham, Qiang Cheng
机构
*
Tsinghua University(清华大学)
;
The University of Manchester(曼彻斯特大学)
;
University of Kentucky(肯塔基大学)
;
Miami University(迈阿密大学)
;
University of Dayton(代顿大学)
;
Institute for Biomedical Informatics, University of Kentucky(肯塔基大学生物医学信息学研究所)
Comments12 pages, 4 figures, 9 tables. v2: adds Adam-mini discussion, learning-rate sweeps with repeated seeds for the AdamW and Adafactor baselines, and a tier ablation; corrects the attribution of the perplexity advantage between momentum and the factored estimator. Code and per-run training logs: https://github.com/nuemaan/skewadam