arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7550 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7550 篇

2606.11657 2026-06-11 cs.LG cs.AI 新提交 82%

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

稀疏探针与模糊物理:连续介质动力学基础模型可解释性挑战的案例研究

Katherine Rosenfeld, Maike Sonnewald

机构 * Gates Foundation(盖茨基金会) UC Davis(加州大学戴维斯分校)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本研究通过稀疏自编码器探针分析连续介质动力学基础模型Walrus的内部机制,发现其内部特征与物理分解不完全一致,并存在输出级偏差,揭示了科学基础模型可解释性的关键挑战。

Comments 8 pages, 5 figures

Journal ref ICLR 2026 Workshop on Foundation Models for Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16327 2026-04-27 cs.AI cs.LG 82%

On the Power of Foundation Models

基础模型的威力

Yang Yuan

机构 * IIIS, Tsinghua University(清华大学信息学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Qi Zhi Institute(上海启智研究院)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG;LLM(comments)

AI总结 本文通过范畴论探讨基础模型在提示学习和微调中的能力限制及泛化理论,提出新的泛化定理。

Comments ICML'23. This version polished paper with the help of LLM, fixed a few notational issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08931 2025-10-13 cs.AI cs.LG 82%

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation

Ashish Kattamuri, Harshwardhan Fartale, Arpita Vats, Rahul Raja, Ishita Prasad

机构 * Proofpoint Indian Institute of Science(印度科学研究院) Linkedin Meta FAIR

专题命中 知识编辑与模型理解 :LLM(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05801 2025-10-07 cs.LG cs.AI 82%

time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models

Debdeep Sanyal, Aaryan Nagpal, Dhruv Kumar, Murari Mandal, Saurabh Deshpande

机构 * Birla AI Labs, Office of Ananya Birla(Birla AI实验室,Ananya Birla办公室) BITS Pilani KIIT Bhubaneswar(KIIT巴尔班格斯)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG

Journal ref NeurIPS 2025 Workshop on Recent Advances in Time Series Foundation Models (BERT2S)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13779 2025-05-21 cs.CL cs.LG 82%

The Mystery of the Pathological Path-star Task for Language Models

Arvid Frydenlund

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.LG

Comments EMNLP 2024 Main at https://aclanthology.org/2024.emnlp-main.695/ See 'Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More' for a follow-up work

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18499 2025-02-27 cs.SE cs.AI cs.CL 82%

Mechanistic Understanding of Language Models in Syntactic Code Completion

Samuel Miller, Daking Rai, Ziyu Yao

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI;foundation model(comments)

Comments 10 pages, 4 figures, accepted to the AAAI 2025 Workshop on Towards Knowledgeable Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14097 2024-12-19 cs.LG cs.AI cs.CV 82%

Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts

Jihye Choi, Jayaram Raghuram, Yixuan Li, Somesh Jha

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG

Comments The preliminary version of the work appeared in the ICML 2024 Workshop on Foundation Models in the Wild

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08065 2024-07-12 cs.RO cs.AI cs.LG 82%

Towards Interpretable Foundation Models of Robot Behavior: A Task Specific Policy Generation Approach

Isaac Sheidlower, Reuben Aronson, Elaine Schaertl Short

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.AI、cs.LG

Comments Short Paper accepted to RLC 2024 Workshop on Training Agents with Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17574 2026-08-20 cs.AI eess.SY 交叉投稿 81%

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

演化不确定性下的风险量化:面向安全序贯决策的信念依赖鲁棒性

Deep Kumar Ganguly, Jan Kretinsky

机构 * Technical University of Munich(慕尼黑工业大学) Masaryk University(马萨里克大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 本文提出RATTL算法,将智能体谨慎程度与认知不确定性绑定,证明其规划问题适定且满足安全三明治性质,可保障LLM等智能体在不确定性下的运行时安全。

Comments Accepted for presentation at the IJCAI-ECAI 2026 RobustifAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01367 2026-08-20 cs.CL stat.ML 版本更新 81%

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

MMD-Flagger:利用最大均值差异检测幻觉

Kensuke Mitsuzawa, Damien Garreau

机构 * Université Côte d’Azur, CNRS, LJAD, France(法国埃克塞特大学、法国国家科学研究中心、LJAD研究所) Center for Artificial Intelligence and Data Science (CAIDAS)(人工智能与数据科学中心) Julius-Maximilans-Universität Würzburg, Germany(德国维尔茨堡约利乌斯-马克斯米利安大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对智能体AI系统中大型语言模型的幻觉问题,提出基于最大均值差异(MMD)的MMD-Flagger方法,通过监控不同解码温度下的输出轨迹检测幻觉,在MUCH基准上用Llama-3、Gemma-3模型完成评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17524 2026-08-19 cs.LG 新提交 81%

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

通过可解释强化学习(XRL)方法对修复智能体Bug的帮助程度来评估XRL方法

Ram Rachum, Yotam Amitai, Bálint Gyevnár, Reuth Mirsky, Cameron Allen

机构 * University of California, Berkeley(加州大学伯克利分校) Tufts University(塔夫茨大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出EvalXRL基准,通过大型语言模型编码智能体结合XRL方法诊断修复RL智能体故障的效果,实现对多种XRL方法的首次闭环直接对比评估。

Journal ref Proceedings of the Workshop on Explainable Artificial Intelligence (XAI) at IJCAI-ECAI 2026, Bremen, Germany

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14707 2026-08-18 cs.AI 新提交 81%

Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems

分层多智能体系统中基于语义不确定性的协调策略

John Knowlton, Aritra Guha, Risto Miikkulainen

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) AT&T Chief Data Office(AT&T首席数据办公室) Cognizant AI Lab(高知特人工智能实验室)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出HASSUM框架,利用语义熵和密度估计不确定性实现多智能体自适应协调,在StrategyQA等基准上验证其可提升复杂推理任务的可靠性。

Comments 17 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12771 2026-08-14 cs.SE cs.AI 新提交 81%

Memorization Diagnostics for Code LLMs Should be Scale-Aware

代码大语言模型的记忆诊断应考虑规模因素

Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对代码大语言模型记忆诊断的规模敏感性问题,通过分离表征负载与记忆,发现现有探测技术在大规模模型上失效,提出未来评估需解耦二者的方法。

Comments 26 pages, 6 figures, 6 tables. Under review at EMSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23300 2026-08-13 q-fin.PM cs.AI cs.MA q-fin.ST 版本更新 81%

Designing Agentic AI-Based Screening for Portfolio Investment

基于代理AI的资产组合投资筛选设计

Mehmet Caner, Agostino Capponi, Nathan Sun, Jonathan Y. Tan

机构 * North Carolina State University, Nelson Hall, Department of Economics(北卡罗来纳州立大学,Nelson Hall,经济学系) Columbia University. Department of Industrial Engineering and Operations Research and Columbia Business School(哥伦比亚大学,工业工程与运筹学系及哥伦比亚商学院) Columbia University. Department of Industrial Engineering and Operations Research(哥伦比亚大学,工业工程与运筹学系)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于代理AI的资产组合投资筛选框架,通过两个专用代理筛选优质公司并生成买卖信号,结合高维精度矩阵估计确定最优权重,实验证明其在S&P 500数据上优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10434 2026-08-12 cs.AI 新提交 81%

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

无人机入侵检测中对话式与仪表盘式可解释人工智能:操作员信任与依赖的实证研究

Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong

机构 * Phenikaa University(菲卡大学) Phenikaa School of Computing(菲卡计算机学院) MobiFone Corporation(MobiFone集团) MobiFone HighTech Center(MobiFone高科技中心) National Economics University(国民经济大学) College of Technology(技术学院)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究对比对话式与仪表盘式XAI界面对无人机入侵检测操作员信任和依赖的影响,发现对话式界面可用性更高但易引发过度依赖,为未来XAI系统设计提供启示。

Comments 12 pages, 3 figures, EIDT conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09248 2026-08-12 cs.AI 版本更新 81%

Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution

Emotion2Skill:用于自适应技能选择与演化的模型内部情绪信号

Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu, Chen Zhang, Yudong Zhang

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 Emotion2Skill框架提取LLM内部情绪向量,将其融入技能选择与演化,在WebShop、ALFWorld上用Qwen3模型实现显著性能提升,验证了情绪表征作为智能体决策信号的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00422 2026-08-12 cs.AI 版本更新 81%

TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs

TrAC:用于大语言模型高效不确定性量化的轨迹条件答案一致性

Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang

机构 * Florida State University(佛罗里达州立大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对大语言模型推理轨迹与答案不一致的问题,提出TrAC框架,结合主动与被动信号实现高效不确定性量化,在数学推理基准上提升了AUROC并降低了AURC。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00143 2026-08-11 cs.CR cs.AI 版本更新 81%

Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity

基于Atomic Red Team技术的符号化攻击链生成:谓词表示粒度的实证研究

Ramya Varunsegar

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究对比九类与五类谓词粒度的AALM,发现粒度对攻击链有效性影响小,81.3%结果一致,高粒度仅提升规划论证的内部结构分辨率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29190 2026-08-03 cs.AI 新提交 81%

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

CAGE:面向使用工具的智能体的带类型返回不确定性的可验证授权

Blaise Delattre, Cong Wang, Yang Cao

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.AI

AI总结 该研究针对使用工具的LLM智能体的授权问题,提出CAGE方法直接验证带绑定错误和数值漂移的联合邻域,可消除点式网关的预算内错误授权,保留部分自主决策,适配多种场景。

Comments Code: https://github.com/tdsai-lab/cage-agent-authorization

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24348 2026-07-28 cs.CR cs.AI 新提交 81%

DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense

深度信念:用于多阶段APT防御中忠实事件报告的基于证据的大语言模型

Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 针对多阶段APT防御中事件报告难检测解释、输出难理解及大语言模型内容缺乏依据等问题,提出深度信念框架,通过集成多种技术确保生成陈述有依据,实验表明该框架能提升多项指标,为安全运营中心提供可靠报告。

Comments This paper has been submitted to the IEEE International Conference on Network and Service Management (CNSM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08992 2026-07-28 cs.AI 版本更新 81%

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

重新思考用于大语言模型的前景理论:在认知不确定性下决策的不稳定性

Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song

机构 * Hong Kong University of Science and Technology(香港科技大学) Huazhong University of Science and Technology(华中科技大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文重新审视前景理论在大语言模型中的应用,发现其在认知不确定性下的决策稳定性存在问题,指出前景理论可能无法可靠地处理此类不确定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15607 2026-07-20 cs.LG stat.ML 新提交 81%

ASK-NN: An Asymmetric Nearest-Neighbor Test that detects Distribution Drifts in Natural Language

ASK-NN:一种检测自然语言中分布漂移的非对称最近邻测试

Sergey Zakharov, Rodion Oblovatny, Alexey Zaytsev

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.LG

AI总结 研究自然语言中分布漂移检测问题,提出基于有向k近邻图的非对称双样本测试ASK-NN,其计算高效易实现,在合成基准、人工文本及LLM幻觉检测方面与核和图基线相比具有竞争力。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08583 2026-07-16 cs.CL 版本更新 81%

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

来源或未发生:一种用于引用伪造检测的多智能体框架

Mingzhe Li, Zhiqiang Lin, Shiqing Ma

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) The Ohio State University(俄亥俄州立大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出CiteTracer多智能体框架,通过12种分类体系检测引用伪造,利用PDF和BibTeX提取结构化引用,结合缓存查找、URL获取等方法验证,实现97.1%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24967 2026-07-07 cs.AI 版本更新 81%

The Anatomy of Uncertainty in LLMs

大语言模型中不确定性的解剖结构

Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy

机构 * School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种将大语言模型不确定性分解为输入模糊性、知识缺口和解码随机性的框架,通过实验展示不同模型规模和任务中各组件的重要性,以提升模型可靠性与可信度。

Comments Accepted at the ICBINB Workshop, ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01773 2026-07-03 cs.AI 新提交 81%

Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

通过检索基础的形式概念分析实现可验证的知识扩展

Yujin Yang, Heejung Lee

机构 * Hanyang University(汉阳大学)

专题命中 知识编辑与模型理解 :SLM(abstract,abstract_cn);language model(abstract);small language model(abstract);分类 cs.AI

AI总结 提出一种检索增强的小语言模型框架,利用形式概念分析作为符号验证循环,通过种子属性和检索验证实现知识扩展,在罕见共济失调数据集上评估了关系F1和蕴含F1。

Comments 8 pages, 2 figures, Accepted to the 8th epiDAMIK ACM SIGKDD International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00012 2026-07-02 cs.IR cs.AI 新提交 81%

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption

PRA-RAG:检索增强生成中针对检索污染的鲁棒聚合方法

Xue Tan, Yi Zheng, Chang Huo, Yunruo Zhang, Yu Liu, Hao Luan, Zhuyang Yu, Xiaoyan Sun, Ping Chen, Jun Dai

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出PRA-RAG算法,通过采样检索文本组合并利用嵌入空间几何结构选择鲁棒子集,提供理论鲁棒性保证,在多个基准上将攻击成功率降至1%同时保持71%准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30973 2026-07-01 cs.CL 新提交 81%

From Propositional to Perceptual Asymmetry: Extending Frictive Policy Optimization to Asymmetric Partial Information Dialogue

从命题不对称到感知不对称:将摩擦策略优化扩展到非对称部分信息对话

Yifan Zhu, Kyeongmin Rim, James Pustejovsky

机构 * Brandeis University(布兰迪斯大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 将摩擦策略优化从共享感知场景扩展到感知不对称场景,通过跨语料库分析和LLM探测验证,发现摩擦函数仅在参与者信息视野内有效,并提出标注改进。

Comments 11 pages. To appear in Proceedings of SIGDIAL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29049 2026-06-30 cs.LG 81%

MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment

MOSAIC: 通过层次语义对齐编排协作知识追踪

Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou, Yu Ma

机构 * Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校) School of Information Systems and Management, Carnegie Mellon University(信息系统与管理学院,卡内基梅隆大学) Department of Mathematics, University of Southern California(数学系,南加州大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Computer Science Department, UC San Diego(计算机科学系,UCSD)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);分类 cs.LG

AI总结 提出MOSAIC框架,利用冻结LLM生成动态嵌入和层次预测提示,结合跨粒度一致性目标,在协作知识追踪中实现多粒度掌握估计,在多个数据集上取得SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28770 2026-06-30 cs.AI 81%

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

LLMs人格的机械论分析:通过潜在特征干预引导人格

David Courtis, Ting Hu

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出一种机械可解释性方法,通过稀疏自编码器和对比激活分析识别残差流中的潜在方向,并施加加法干预向量来增强目标OCEAN人格特质,同时保持语言建模性能。

Comments Written in 2024; submitted to arXiv 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24050 2026-06-30 cs.LG stat.ML 81%

A Mechanistic Study of Transformers Training Dynamics

Transformer训练动态的机制研究

Ambroise Odonnat, Wassim Bouaziz, Vivien Cabannes

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);foundation model(abstract);pretraining(abstract)

AI总结 本文通过可控实验研究Transformer训练动态,发现梯度下降可实现聚类头解决稀疏模块加法任务,并揭示训练过程中两阶段学习及归一化层高曲率导致的损失尖峰现象。

Comments Accepted at ICML 2026 Mechanistic Interpretability workshop

详情

展开后加载摘要…

URL PDF HTML 收藏