arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-07-03 至 2026-07-03 共收录 21 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 21 篇

2607.01235 2026-07-03 cs.CL cs.AI cs.SE 新提交 92%

TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

TokenScope: 面向大型语言模型中代码任务的令牌级可解释性与解释性

Amirreza Esmaeili, Fatemeh Fard

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 提出TokenScope工具,通过令牌级指标、注意力模式和抽象语法树聚合,实现解码时信号与结构分析的统一,支持交互式令牌替换和反事实分支,以系统研究LLM在代码生成中的行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19775 2026-07-03 cs.AI cs.CL cs.ET cs.MA cs.RO 版本更新 92%

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

从行为到理解:LLM代理中时间概念的符合可解释性

Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy

机构 * Georgia State University(佐治亚州立大学) Computer Science Lab, SRI(SRI计算机科学实验室) Army Cyber Institute(陆军网络学院) United States Military Academy(美国军事学院) University of Florida(佛罗里达大学)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过符合视角解释LLM代理时间概念演变的框架,结合逐步奖励建模与符合预测,识别时间概念的潜在方向,实验表明这些概念可线性分离,为LLM代理提供可靠故障检测和干预方法。

Comments Accepted at the Mechanistic Interpretability Workshop, 43rd International Conference on Machine Learning, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02052 2026-07-03 cs.SE 新提交 92%

Mitigating Package Hallucinations in Large Language Models via Model Editing

通过模型编辑缓解大型语言模型中的包幻觉

Shuhan Liu, Yukai Zhao, Xing Hu, Kui Liu, Xiaohu Yang, Xin Xia

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract)

AI总结 提出BOUND框架,通过定位与包幻觉相关的模块并应用边界感知的LoRA适配器编辑,有效降低LLM在包推荐、代码生成等任务中的包幻觉率,同时保持有效包推荐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02455 2026-07-03 cs.HC 新提交 91%

When Do LLM Personas Support Visualization Design? A Cross-Model Study of Color Assignment and Chart Choice

LLM人格何时支持可视化设计?跨模型研究颜色分配与图表选择

Shahreen Salim, Klaus Mueller

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本研究通过跨模型实验(GPT-4o-mini、GPT-4.1-mini、GPT-5-mini)和43个大五人格剖面,发现LLM人格在颜色分配任务中表现不稳定,而在图表选择任务中任务上下文比人格影响更大,表明LLM人格仅适合作为探索性工具而非人类替代。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02182 2026-07-03 cs.LG cs.CL 新提交 91%

Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

贝叶斯稀疏低秩自适应用于大语言模型不确定性估计

Jijie Zhang, Zhe Ren, Quan Zhang, Dandan Guo

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Michigan State University(密歇根州立大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.CL、cs.LG

AI总结 提出DALorRA,一种变分贝叶斯稀疏框架,通过随机掩码低秩适应中的秩维度实现模型容量正则化和校准,在不牺牲推理精度下提升LLM校准性能。

Comments Preprint. 16 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01457 2026-07-03 cs.CL cs.AI 新提交 89%

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

接地优化:一种减少自动个人文档重写中LLM幻觉的分层工程框架

Shashank Indukuri, Adarsh Agrawal

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出五层接地优化框架,通过时间验证、污染检测、结构不变性、提示接地和评估器,将简历重写中的幻觉率降至0.04-0.24。

Comments 13 pages, 1 figure. Equal contribution by both authors. Code and data: https://github.com/shashank-indukuri/grounded-optimization

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06324 2026-07-03 cs.SE cs.MA 版本更新 87%

From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws

从失败轨迹到可靠的LLM代理:诊断与修复框架缺陷

Mengzhuo Chen, Junjie Wang, Zhe Liu, Yawen Wang, Haiming Zheng, Qing Wang

专题命中 知识编辑与模型理解 :LLM(title,title_cn)

AI总结 提出HarnessFix框架,通过轨迹引导诊断代理失败并修复执行框架,在多个基准上提升15.2%-50.0%性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01901 2026-07-03 cs.LG cs.AI cs.MM 新提交 82%

SABER: A Semantic-Aligned Brain Network Analysis Framework via Multi-scale Hypergraphs

SABER: 一种基于多尺度超图的语义对齐脑网络分析框架

Yidan Xu, Xiangmin Han, Rundong Xue, Huihui Ye

机构 * Hangzhou Dianzi University(杭州电子科技大学) Tsinghua University(清华大学) Xi’an Jiaotong University(西安交通大学)

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出SABER框架,通过多尺度超图建模功能子网络与高阶依赖,并引入决策级语义对齐机制,将大语言模型语义直接融入预测,在ABIDE和ADHD-200数据集上取得最优性能,尤其在小样本场景下。

Comments Accepted to IEEE International Conference on Multimedia and Expo (ICME) 2026;

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01773 2026-07-03 cs.AI 新提交 81%

Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

通过检索基础的形式概念分析实现可验证的知识扩展

Yujin Yang, Heejung Lee

机构 * Hanyang University(汉阳大学)

专题命中 知识编辑与模型理解 :SLM(abstract,abstract_cn);language model(abstract);small language model(abstract);分类 cs.AI

AI总结 提出一种检索增强的小语言模型框架,利用形式概念分析作为符号验证循环,通过种子属性和检索验证实现知识扩展,在罕见共济失调数据集上评估了关系F1和蕴含F1。

Comments 8 pages, 2 figures, Accepted to the 8th epiDAMIK ACM SIGKDD International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02217 2026-07-03 q-bio.GN 新提交 80%

Affinage: genome-scale mechanistic gene annotation from the published literature

Affinage: 基于已发表文献的全基因组尺度机制性基因注释

Matteo Di Bernardo, Iain M. Cheeseman

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出Affinage,一个利用大语言模型从原始文献中提取直接实验证据并推理,为人类蛋白质编码基因生成可复用的结构化机制注释的流程,在19,293个基因上优于UniProt。

Comments 12 pages, 6 figures. Accepted to the Workshop on Generative and Agentic AI for Biology at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29341 2026-07-03 cs.IR 版本更新 80%

Monosemanticity in Recommender Systems

推荐系统中的单语义性

Yagel Alfasi, Eden Rzezak, Eadan Schechter

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract)

AI总结 研究矩阵分解嵌入的层次稀疏表示,用Matryoshka稀疏自编码器提取可解释特征,通过元数据对齐和LLM标注验证语义一致性,并干预性别相关神经元。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20237 2026-07-03 cs.CL 79%

Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models

探讨在微调语言模型中回话用语和填充词的表示

Yu Wang, Leyi Lao, Langchu Huang, Gabriel Skantze, Yang Xu, Hendrik Buschmeier

机构 * Bielefeld University(比勒菲尔德大学) Southern University of Science and Technology(南方科技大学) KTH Royal Institute of Technology(皇家理工学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL

AI总结 本文通过三种微调策略研究回话用语和填充词在语言模型中的表示,发现微调能提升模型区分语义差异的能力,使生成的对话更接近人类语言。

Comments Accepted at ACL 2026 main

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers, ACL 2026), pp. 5319-5348

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29466 2026-07-03 cs.LG cs.AI cs.CL 版本更新 75%

An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms

基于梯度范数的高效不确定性量化方法

Nils Grünefeld, Jes Frellsen, Christian Hardmeier

机构 * IT University of Copenhagen(哥本哈根信息技术大学) Pioneer Centre for Artificial Intelligence(人工智能先锋中心) Technical University of Denmark(丹麦技术大学)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种轻量级方法,通过梯度范数和参数协方差的近似,实现对神经网络预测不确定性的高效量化,验证了其在不同基准测试中的有效性。

Comments ProbML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02396 2026-07-03 cs.AI cs.LG 新提交 73%

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

通过RFM-AGOP的快速多维拒绝子空间

Thomas Winninger

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出利用RFM-AGOP算法快速提取大型语言模型中的多维拒绝子空间,在推理和非推理模型上均实现秒级识别,且性能优于现有方法。

Comments Accepted to the Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08236 2026-07-03 cs.CL cs.LG 新提交 73%

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

共享语义,不同机制:通过对齐语义与机制的无监督特征发现

Hyunjin Cho, Youngji Roh, Jaehyung Kim

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 提出一种无监督方法,通过语义嵌入和归因签名聚类模型续写,发现隐藏的机制模式,补充电路分析。

Comments 40 pages; accepted as an ICML 2026 Spotlight; project page: https://merenova.github.io/distribution-level-feature-discovery/

Journal ref ICML 2026 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02494 2026-07-03 cs.CV cs.CL 新提交 70%

Towards Robustness against Typographic Attack with Training-free Concept Localization

基于训练无关的概念定位实现对抗印刷体攻击的鲁棒性

Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 知识编辑与模型理解 :language model(abstract);pretraining(abstract);分类 cs.CL

AI总结 提出一种无需训练的可解释性方法,通过分析注意力头对词汇和语义的编码差异,定位并干预ViT中的词汇偏置电路,从而在不额外训练的情况下显著提升对印刷体攻击的鲁棒性。

Comments 15 pages main text, provisionally accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02186 2026-07-03 cs.AI 新提交 70%

UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development

UA-ChatDev: 不确定性感知的多智能体协作实现可靠软件开发

Temitayo Olamilekan Ogunsusi, Lijun Qian, Xishuang Dong

机构 * Department of Electrical Computer Engineering Prairie View A\&M University, Prairie View, TX 77446, USA Email

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出UA-ChatDev框架,通过不确定性量化与阶段感知阈值校准,在软件开发多智能体协作中抑制幻觉传播,提升代码执行可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01336 2026-07-03 quant-ph cond-mat.dis-nn cond-mat.str-el cs.AI cs.LG 新提交 62%

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

神经量子态的机械可解释性与因果特征引导:基于稀疏自编码器

Zihao Qi, Christopher Earls

机构 * Department of Physics, Cornell University(康奈尔大学物理系) Center for Applied Mathematics, Cornell University(康奈尔大学应用数学中心)

专题命中 知识编辑与模型理解 :post-training(abstract);分类 cs.AI、cs.LG

AI总结 利用稀疏自编码器从神经量子态内部激活中无监督提取物理可观测量相关特征,并通过后训练干预单一特征因果地平滑引导对应可观测量,揭示其可解释内部表示。

Comments 15 pages, 7 figures. Comments welcome!

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01303 2026-07-03 cs.CV cs.AI 新提交 57%

CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection

CPG-PAD: 概念引导提示的呈现攻击检测

Haoyuan Zhang, Xiangyu Zhu, Li Gao, Ajian Liu, Siran Peng, Zhen Lei

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) China Mobile Financial Technology Co., Ltd.(中移动金融科技有限公司) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心) School of Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院计算机科学与工程学院)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.AI

AI总结 提出概念引导提示的呈现攻击检测框架,通过可解释AI发现攻击相关视觉概念并注入提示空间,提升跨域泛化能力,在多个基准数据集上达到最优。

Comments Accepted by IEEE Transactions on Information Forensics & Security (TIFS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07820 2026-07-03 cs.CL 57%

Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests

参考游戏作为模型不确定性与澄清请求对齐的测试平台

Manar Ali, Judith Sieker, Sina Zarrieß, Hendrik Buschmeier

机构 * Digital Linguistics Lab(数字语言实验室) Computational Linguistics Group(计算语言学小组)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL

AI总结 本文通过参考游戏测试语言模型在不确定性识别与澄清请求表达上的能力,发现模型在简单任务中难以准确识别自身不确定性并转化为澄清行为。

Comments Accepted at GEM@ACL 2026, the 5th Generation, Evaluation & Metrics Workshop

Journal ref Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM 2026), pp. 990-998

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01630 2026-07-03 cs.CV 新提交 50%

DRDN: Decoupled Representation Dynamic Network for From-Scratch ViT Class-Incremental Learning

DRDN: 用于从头训练ViT类增量学习的解耦表示动态网络

Bingchen Huang, Yifu Chen, Zhiling Wang, Yuanchao Du

机构 * Meituan Vision AI Department(美团视觉AI部门)

专题命中 知识编辑与模型理解 :pretraining(abstract)

AI总结 针对类增量学习中共享表示退化与任务间干扰问题,提出解耦表示动态网络(DRDN),通过掩码图像建模保持骨干网络通用视觉结构,并采用层次化任务令牌扩展减少跨任务干扰,在从头训练的ViT上取得显著提升。

Comments 10 pages, IEEEtran journal format. Preprint submitted to IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏