arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 7552 信号源:cs.CL, cs.AI, cs.LG

1. 知识编辑与模型理解 7552 篇

2603.04972 2026-03-06 cs.LG cs.CL 76%

Functionality-Oriented LLM Merging on the Fisher--Rao Manifold

面向功能的LLM在Fisher-Rao流形上的融合

Jiayu Wang, Zuojun Ye, Wenpeng Yin

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.LG

AI总结 本文提出在Fisher-Rao流形上计算加权Karcher均值,以实现更稳定的LLM融合,提升模型性能。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06129 2026-02-09 cs.LG cs.AI 76%

Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction

城市时空基础模型用于气候韧性住房:扩散变换器的扩展用于灾害风险预测

Olaf Yunus Laitinen Imanov, Derya Umut Kulali, Taner Yilmaz

机构 * Technical University of Denmark(技术大学) Eskisehir Technical University(埃斯基谢普大学) Afyon Kocatepe University(阿夫yon卡奥塔佩大学)

专题命中 知识编辑与模型理解 :foundation model(title);分类 cs.AI、cs.LG

AI总结 本文提出Skjold-DiT模型,通过整合时空城市数据预测建筑气候风险,并结合交通网络结构提升灾害响应能力。

Comments 10 pages, 5 figures. Submitted to IEEE Transactions on Intelligent Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16766 2026-01-26 cs.CL cs.AI 76%

Do LLM hallucination detectors suffer from low-resource effect?

大型语言模型的幻觉检测器是否受到低资源效应影响?

Debtanu Datta, Mohan Kishore Chilukuri, Yash Kumar, Saptarshi Ghosh, Muhammad Bilal Zafar

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI

AI总结 研究发现,幻觉检测器在低资源语言中表现较任务本身更稳健,可能因内部机制编码了不确定性信号。

Comments Accepted at EACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24793 2025-12-18 cs.SD cs.AI cs.LG eess.AS 76%

Sparse Autoencoders Make Audio Foundation Models more Explainable

稀疏自编码器使音频基础模型更加可解释

Théo Mariotte, Martin Lebourdais, Antonio Almudévar, Marie Tahon, Alfonso Ortega, Nicolas Dugué

机构 * LIUM, Le Mans Université(利姆大学LIUM) VivoLab, I3A, University of Zaragoza(瓦沃拉实验室、I3A、萨拉戈萨大学)

专题命中 知识编辑与模型理解 :foundation model(title);分类 cs.AI、cs.LG

AI总结 本研究利用稀疏自编码器分析音频预训练模型的隐藏表示,揭示其在歌唱技巧分类中的可解释性提升及对自监督学习系统的作用。

Comments 5 pages, 5 figures, 1 table, submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12699 2025-10-15 cs.CL cs.AI 76%

Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations

Sunny Yu, Ahmad Jabbar, Robert Hawkins, Dan Jurafsky, Myra Cheng

机构 * Department of Computer Science, Stanford University(计算机科学系,斯坦福大学) Department of Linguistics, Stanford University(语言学系,斯坦福大学)

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15038 2025-07-31 cs.CL cs.AI 76%

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

Haiyan Zhao, Xuansheng Wu, Fan Yang, Bo Shen, Ninghao Liu, Mengnan Du

机构 * New Jersey Institute of Technology(新泽西理工学院) University of Georgia(佐治亚大学) Wake Forest University(威克森林大学)

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10015 2025-07-18 cs.CV cs.AI cs.LG 76%

(Almost) Free Modality Stitching of Foundation Models

Jaisidh Singh, Diganta Misra, Boris Knyazev, Antonio Orvieto

机构 * University of Tübingen(图宾根大学) Zuse School ELIZA(Zuse学校ELIZA) ELLIS Institute Tübingen(图宾根ELLIS研究所) MPI-IS Tübingen(图宾根MPI-IS研究所) SAIT AI Lab Montréal(蒙特利尔SAIT人工智能实验室) Tübingen AI Center(图宾根人工智能中心)

专题命中 知识编辑与模型理解 :foundation model(title);分类 cs.AI、cs.LG

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08852 2025-06-18 cs.AI cs.LG 76%

A Unified Framework for Next-Gen Urban Forecasting via LLM-driven Dependency Retrieval and GeoTransformer

Yuhao Jia, Zile Wu, Shengao Yi, Yifei Sun, Xiao Huang

机构 * Emory University, University of Pennsylvania(埃默里大学和宾夕法尼亚大学) University of Pennsylvania(宾夕法尼亚大学) Emory University(埃默里大学)

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02481 2025-06-04 cs.CL cs.AI 76%

Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths

Inderjeet Nair, Lu Wang

机构 * University of Michigan(密歇根大学)

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20487 2025-05-28 cs.CL cs.AI 76%

InFact: Informativeness Alignment for Improved LLM Factuality

Roi Cohen, Russa Biswas, Gerard de Melo

机构 * Hasso Plattner Institute University of Potsdam(霍普夫纳研究所波茨坦大学) Dept. of Computer Science Aalborg University(计算机科学系奥尔堡大学)

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11333 2025-05-27 cs.CL cs.AI 76%

Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models

Xiaochen Zhu, Georgi Karadzhov, Chenxi Whitehouse, Andreas Vlachos

机构 * University of Cambridge(剑桥大学)

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

Comments 9 pages (main body), 3 figures (main body), ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19889 2024-10-29 cs.CL cs.LG 76%

Ensembling Finetuned Language Models for Text Classification

Sebastian Pineda Arango, Maciej Janowski, Lennart Purucker, Arber Zela, Frank Hutter, Josif Grabocka

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

Comments Workshop on Fine-Tuning in Modern Machine Learning @ NeurIPS 2024. arXiv admin note: text overlap with arXiv:2410.04520

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11498 2024-09-19 cs.SD cs.AI cs.CL eess.AS 76%

Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning

Ilaria Manco, Justin Salamon, Oriol Nieto

专题命中 知识编辑与模型理解 :LLM(title);分类 cs.CL、cs.AI

Comments To appear in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05976 2024-08-13 cs.LG cs.CL 76%

Global-to-Local Support Spectrums for Language Model Explainability

Lucas Agussurja, Xinyang Lu, Bryan Kian Hsiang Low

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06730 2024-07-11 cs.CL cs.AI cs.HC 76%

Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Maarten Sap

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

Comments ACL 2024 (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09145 2024-06-12 cs.CL cs.AI 76%

ToNER: Type-oriented Named Entity Recognition with Generative Language Model

Guochao Jiang, Ziqin Luo, Yuchen Shi, Dixuan Wang, Jiaqing Liang, Deqing Yang

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06795 2024-06-04 cs.CL cs.LG 76%

Robust Infidelity: When Faithfulness Measures on Masked Language Models Are Misleading

Evan Crothers, Herna Viktor, Nathalie Japkowicz

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15823 2024-01-23 cs.CL cs.AI 76%

Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment

Ahmed ElBakry, Mohamed Gabr, Muhammad ElNokrashy, Badr AlKhamissi

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.AI

Comments Proceedings of ArabicNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06427 2023-11-14 cs.CL cs.LG 76%

ChatGPT Prompting Cannot Estimate Predictive Uncertainty in High-Resource Languages

Martino Pelucchi, Matias Valdenegro-Toro

专题命中 知识编辑与模型理解 :prompting(title);分类 cs.CL、cs.LG

Comments 14 pages, 4 figures, with appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15837 2022-11-30 cs.LG cs.AI cs.CV cs.GT 76%

Survey on Self-Supervised Multimodal Representation Learning and Foundation Models

Sushil Thapa

专题命中 知识编辑与模型理解 :foundation model(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.15222 2021-03-30 cs.CL cs.LG q-bio.BM 76%

BERTology Meets Biology: Interpreting Attention in Protein Language Models

Jesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

Comments To appear in ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05295 2020-05-14 cs.CL cs.LG stat.ML 76%

Language Models Are An Effective Patient Representation Learning Technique For Electronic Health Record Data

Ethan Steinberg, Ken Jung, Jason A. Fries, Conor K. Corbin, Stephen R. Pfohl, Nigam H. Shah

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.01817 2019-08-07 cs.CL cs.LG stat.ML 76%

Sparsity Emerges Naturally in Neural Language Models

Naomi Saphra, Adam Lopez

专题命中 知识编辑与模型理解 :language model(title);分类 cs.CL、cs.LG

Comments Published in the ICML 2019 Workshop on Identifying and Understanding Deep Learning Phenomena: https://openreview.net/forum?id=H1ets1h56E

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25379 2026-03-27 cs.AI cs.HC 76%

Does Structured Intent Representation Generalize? A Cross-Language, Cross-Model Empirical Study of 5W3H Prompting

结构化意图表示是否具有泛化性?一项跨语言、跨模型的5W3H提示经验研究

Peng Gang

机构 * Huizhou Lateni AI Technology Co., Ltd.(惠州拉特尼人工智能科技有限公司) Huizhou University(惠州大学)

专题命中 知识编辑与模型理解 :prompting(title,comments);分类 cs.AI

AI总结 本文通过跨语言和跨模型的实验,探讨结构化意图表示的泛化能力,发现AI辅助写作工具生成的5W3H提示在目标对齐上与手动创建的提示无显著差异,且能减少跨模型输出方差。

Comments 28 pages, figures, tables, and appendix. Follow-up empirical study extending prior work on PPS and 5W3H structured prompting to cross-language, cross-model, and AI-assisted authoring settings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19611 2026-08-21 cs.CL cs.AI cs.LG 新提交 75%

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

快速分叉:高效估计文本生成中的不确定性动态

Eric Bigelow, Amir Zur, Satchel Grant, Tal Haklay, Can Rager, Owen Lewis, Thomas McGrath, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 针对文本生成中不确定性动态估计成本高的问题,本研究开发统计模型平滑低采样推理数据以近似高采样数据,提升重采样分析效率,揭示不确定性动态的稳定模式及噪声来源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17099 2026-08-19 cs.HC cs.CY 新提交 75%

Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens

看似合法还不够:通过参与式设计视角审视表征过程中的合成智能体

Aditya Nayak, Aditi Vashistha, Alissa Centivany, Aakash Gautam

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);foundation model(abstract)

AI总结 本文从参与式设计视角,结合三个不同规模的案例,探讨用合成智能体替代人类参与表征过程的合法性与人格关联,指出其风险并提出监督边界。

Comments 13 pages total, 4 figures, accepted to the 9th AAAI Conference on AI, Ethics, and Society (AIES 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09928 2026-08-11 cs.CV cs.AI cs.CL cs.LG 新提交 75%

Multimodal Model Diffing for Feature Discovery and Control

用于特征发现与控制的多模态模型差异分析

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark

机构 * University of Oxford(牛津大学) Microsoft(微软公司)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出MMDiff多模态模型差异分析框架,训练多模态SAEs以识别多模态训练改变的特征,实现特征隔离、检测与控制,在空间、OCR任务及多模态安全攻击评估中展现出良好效果。

Comments Preprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24730 2026-08-06 cs.HC 版本更新 75%

Diamonds in the rough: Transforming SPARCs of imagination into a game concept by leveraging medium sized LLMs

粗糙中的钻石:通过利用中型大语言模型将想象力的SPARCs转化为游戏概念

Julian Geheeb, Farhan Abid Ivan, Daniel Dyrda, Miriam Anschütz, Georg Groh

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文研究了中型大语言模型在早期游戏设计中的应用,通过生成和评估游戏创意,展示了这些模型在提供有用反馈方面的潜力,并指出需要进一步优化提示方法以提高一致性。

Comments 13 pages, 4 figures, 2 tables

Journal ref Proceedings of AI4HGI '25, the First Workshop on Artificial Intelligence for Human-Game Interaction at the 28th European Conference on Artificial Intelligence (ECAI '25), Bologna, October 25-30, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21022 2026-08-04 cs.SE 版本更新 75%

Don't Use a Cannon to Kill a Fly: Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations

轻量级模型编辑用于LLMs纠正过时的API推荐

Guancheng Lin, Xiao Yu, Jacky Keung, Xing Hu, Xin Xia, Alex X. Liu

专题命中 知识编辑与模型理解 :LLM(abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出AdaLoRA-L方法,通过区分通用API层和特定API层,提升LLMs生成最新API的能力。

Comments Accepted to ISSTA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24884 2026-07-30 cs.SE cs.AI cs.CL cs.LG 版本更新 75%

Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation

超越“检索什么”:检索增强代码生成中的不确定性

Chandan Kumar Sah, Li Zhang, Xiaoli Lian

专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 研究存储库级代码生成中异构证据不确定性问题,提出不确定性感知框架OpenCoder,通过估计特定源不确定性来过滤和排序证据以指导相关操作,实验表明其能提高输出正确性,支持将不确定性作为可操作控制信号。

Comments 9 pages, 4 figures. Source code and supporting materials are available at https://github.com/Rocky5502/OpenCoder_V1

详情

展开后加载摘要…

URL PDF HTML 收藏