arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-29 至 2026-04-29 共收录 13 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 13 篇

2604.25011 2026-04-29 cs.CL 93%

Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models

为什么强化学习能够泛化?对大语言模型后训练的特征层面机制研究

Dan Shi, Zhuowen Han, Simon Ostermann, Renren Jin, Josef van Genabith, Deyi Xiong

机构 * TJUNLP Lab, School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院TJUNLP实验室) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Saarland University(萨尔兰大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);post-training(title,abstract);SFT(abstract,abstract_cn)

AI总结 研究通过对比强化学习与监督微调在大语言模型中的表现,揭示其泛化机制,发现强化学习通过持续演变的特征保持基模型表示,而监督微调引入稳定但专用的特征。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10692 2026-04-29 cs.CL 90%

Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

无条件真实性:学习大语言模型的无条件不确定性

Artem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Gleb Kuzmin, Ivan Lazichny, Alexander Panchenko, Preslav Nakov, Timothy Baldwin, Maxim Panov, Artem Shelmanov

机构 * MBZUAI Skoltech(斯克利切夫斯基因工大学) AIRI ETH Zürich(苏黎世联邦理工学院) The University of Melbourne(墨尔本大学) FRC CSC RAS(俄罗斯科学院计算技术研究所)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出通过学习注意力特征来解决大语言模型生成过程中不确定性评分的难题,通过两阶段训练方法提升了选择性生成的效果,优于其他无监督和监督方法。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14427 2026-04-29 cs.CL 89%

Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models

基于令牌密度的不确定性量化方法用于获取大语言模型的真实性

Artem Vazhentsev, Lyudmila Rvanova, Ivan Lazichny, Alexander Panchenko, Maxim Panov, Timothy Baldwin, Artem Shelmanov

机构 * Skoltech(斯克里普切尔技术学院) AIRI(人工智能研究所) MBZUAI(马斯克商学院人工智能研究所) MIPT(莫斯科国立信息技术机械与光学学院) FRC CSC RAS(俄罗斯科学院远东分院计算机科学与控制学院) The University of Melbourne(墨尔本大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文提出一种基于令牌密度的不确定性量化方法,通过提取多层嵌入并计算Mahalanobis距离,提升文本生成和事实核查任务的不确定性评估精度与效率。

Journal ref Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25591 2026-04-29 eess.AS cs.AI cs.CL cs.LG cs.SD 89%

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

穿越不确定性:音频感知大语言模型不确定性估计的实证研究

Chun-Yi Kuan, Wei-Ping Huang, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾大学通讯工程研究所) Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(台湾大学人工智能卓越研究中心)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了音频感知大语言模型的不确定性估计,通过多种方法对比发现语义层面方法在通用音频推理中表现更优,且在可靠性导向任务中效果依赖模型和基准。

Comments Manuscript in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25167 2026-04-29 cs.AI 88%

From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models

从洞察到行动:一种新的框架,用于基于可解释性指导的数据选择在大语言模型中

Ling Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang, Heng Liu, Yangyang Liu, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo

机构 * TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China(天津大学计算机科学与技术学院TJUNLP实验室,中国) Alibaba Group, China(阿里巴巴集团,中国)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出IGDS框架,通过内部任务特征识别和选择共振数据来提升大语言模型性能,实验显示在数学推理任务中数据效率提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25264 2026-04-29 cs.CR cs.SE 86%

MARD: A Multi-Agent Framework for Robust Android Malware Detection

MARD:一种用于鲁棒Android恶意软件检测的多智能体框架

Xueying Zeng, Youquan Xian, Sihao Liu, Xudong Mou, Yanze Li, Lei Cui, Bo Li

专题命中 知识编辑与模型理解 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract)

AI总结 MARD通过结合LLM的语义理解和传统静态分析,提出多智能体框架,有效降低APK深度分析成本,实现高可解释性检测,F1得分达93.46%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23719 2026-04-29 cs.CL cs.AI 84%

AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models

AIPsy-Affect:一种无关键词的临床刺激电池,用于语言模型中情绪的机制可解释性

Michael Keeman

机构 * Keido Labs(Keido实验室)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出AIPsy-Affect,一种480项的临床刺激电池,通过叙事情境引发Plutchik的八种基本情绪,消除关键词混淆,支持情绪机制的可解释性研究。

Comments Dataset paper. 12 pages + appendix, 2 figures. Dataset available at https://huggingface.co/datasets/keidolabs/aipsy-affect. MIT license

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12553 2026-04-29 cs.CL cs.AI 81%

Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility

这仅仅是幻想吗?语言模型表示反映了人类对事件可能性的判断

Michael A. Lepori, Jennifer Hu, Ishita Dasgupta, Roma Patel, Thomas Serre, Ellie Pavlick

机构 * Department of Computer Science(计算机科学系) Department of Cognitive Science(认知科学系) Google(谷歌) DeepMind(深度Mind) Brown University(布朗大学) Johns Hopkins University(约翰霍普金斯大学) Department of Cognitive(认知系) Department of Computer Science & Psychological Sciences(计算机科学与心理学科学系)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过分析语言模型的模态差异向量,揭示其在事件可能性判断上的能力,发现模型随着训练步骤和参数量增加,能更准确地区分模态类别,并与人类判断行为相关联。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25642 2026-04-29 cs.CV cs.AI 79%

Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models

预填充阶段干预用于缓解大视觉-语言模型中的幻觉

Chengsheng Zhang, Chenghao Sun, Xinyan Jiang, Wei Li, Xinmei Tian

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Advanced Research Institute, Chinese Academy of Sciences(上海先进研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出PTI方法,在预填充阶段干预KV缓存以减少幻觉,通过模态感知方向修正错误表示,提升模型可靠性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07802 2026-04-29 cs.CV cs.AI 79%

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

潜在异常知识挖掘:揭示视觉-语言模型中的稀疏敏感神经元

Shaotian Li, Shangze Li, Chuancheng Shi, Wenhua Wu, Yanqiu Wu, Xiaohan Yu, Fei Shen, Tat-Seng Chua

机构 * Macquarie University(麦考瑞大学) Nanjing University of Science and Technology(南京理工大学) The University of Sydney(悉尼大学) National University of Singapore(新加坡国立大学)

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI

AI总结 本文提出LAKE框架,通过挖掘视觉-语言模型中稀疏敏感神经元,实现异常检测的内在可解释性,实验表明其在工业基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21517 2026-04-29 cs.CL cs.AI 62%

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

声音、偏见与指称:语音翻译中性别可解释性研究

Lina Conti, Dennis Fucci, Marco Gaido, Matteo Negri, Guillaume Wisniewski, Luisa Bentivogli

机构 * Fondazione Bruno Kessler(布鲁诺·科塞拉基金会) University of Trento(特伦托大学) Laboratoire de Linguistique Formelle, Université Paris Cité, CNRS(巴黎城市大学语言学实验室,法国国家科学研究中心)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨语音翻译模型如何根据语音特征分配指称词性别,发现模型通过第一人称代词链接性别化术语与说话者,利用频谱分布而非音高集中信息进行性别判断。

Comments Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25902 2026-04-29 cs.CL cs.AI cs.LG 56%

Toward a Functional Geometric Algebra for Natural Language Semantics

迈向自然语言语义的函数几何代数

James Pustejovsky

机构 * Computer Science Department(计算机科学系)

专题命中 知识编辑与模型理解 :分类 cs.CL、cs.AI、cs.LG;language model(comments)

AI总结 本文提出函数几何代数框架,利用克莱因代数提升语义表示的结构组织性,解决传统线性代数在组合语义、类型敏感性和可解释性上的局限。

Comments 43 pages. Keywords: geometric algebra, Clifford algebra, compositional semantics, natural language semantics, type coercion, multivector representations, graded type system, Generative Lexicon, neural language models, distributional semantics

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25102 2026-04-29 cs.CV 50%

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

一个扰动,两种失效模式:通过嵌入引导的字形扰动探测VLM安全性

Ravikumar Balakrishnan, Sanket Mendapara

机构 * Cisco Systems(思科系统)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文通过实证研究揭示多模态嵌入距离对VLM攻击成功率的预测作用,并提出基于嵌入引导的字形扰动方法,验证了可读性与安全对齐的交互影响。

详情

展开后加载摘要…

URL PDF HTML 收藏