arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2401.00396 2024-05-20 cs.CL 74%

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, Kashun Shum, Randy Zhong, Juntong Song, Tong Zhang

专题命中 幻觉与事实性 :trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06764 2024-04-19 cs.AI 74%

GLaM: Fine-Tuning Large Language Models for Domain Knowledge Graph Alignment via Neighborhood Partitioning and Generative Subgraph Encoding

Stefan Dernbach, Khushbu Agarwal, Alejandro Zuniga, Michael Henry, Sutanay Choudhury

专题命中 幻觉与事实性 :alignment(title);分类 cs.AI

Comments Published in AAAI Spring Symposium: AAAI-MAKE 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01928 2023-09-06 cs.RO cs.AI stat.AP 74%

Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners

Allen Z. Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, Zhenjia Xu, Dorsa Sadigh, Andy Zeng, Anirudha Majumdar

专题命中 幻觉与事实性 :alignment(title);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2023, Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04906 2023-04-12 cs.LG cs.CV 74%

Survey on Leveraging Uncertainty Estimation Towards Trustworthy Deep Neural Networks: The Case of Reject Option and Post-training Processing

Mehedi Hasan, Moloud Abdar, Abbas Khosravi, Uwe Aickelin, Pietro Lio', Ibrahim Hossain, Ashikur Rahman, Saeid Nahavandi

专题命中 幻觉与事实性 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13476 2023-02-01 cs.SE cs.LG 74%

An investigation of challenges encountered when specifying training data and runtime monitors for safety critical ML applications

Hans-Martin Heyn, Eric Knauss, Iswarya Malleswaran, Shruthi Dinakaran

专题命中 幻觉与事实性 :safety(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01478 2021-04-26 cond-mat.mtrl-sci cs.CV cs.LG physics.app-ph 74%

Leveraging Uncertainty from Deep Learning for Trustworthy Materials Discovery Workflows

Jize Zhang, Bhavya Kailkhura, T. Yong-Jin Han

专题命中 幻觉与事实性 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04274 2026-06-04 cs.CL cs.CY 73%

Long Live Fine-Tuning: Task-Specific Transformers Outperform Zero-Shot LLMs for Misinformation Response Classification on Reddit

长存微调:在Reddit上,任务特定Transformer在错误信息响应分类中优于零样本LLM

JooYoung Lee, Lin Tian, Angela Brillantes, Adriana-Simona Mihăiţă, Marian-Andrei Rizoiu

机构 * University of Technology Sydney(悉尼技术大学)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);分类 cs.CL、cs.CY

AI总结 通过对比微调模型与零样本LLM在Reddit错误信息评论分类上的表现,发现微调RoBERTa在宏F1分数上显著优于最佳零样本模型,且成本更低,表明任务特定微调在检测隐性错误信息方面仍更可靠。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24286 2026-05-26 cs.LG cs.CL 73%

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

忠实性作为信息流:评估与训练忠实的链式思维推理

Jinghan Jia, Joe Benton, Eric Easley

机构 * Dept. CSE, Michigan State University(密歇根州立大学计算机科学系) Anthropic

专题命中 幻觉与事实性 :safety(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 通过信息流视角提出基于充分性、完整性和必要性的框架,结合熵、掩码KL和梯度诊断评估链式思维忠实性,并引入更新时干预(如注意力掩码、反向梯度掩码等)训练更忠实的推理模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20591 2026-05-21 cs.CL cs.CY 73%

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

有害吗?网络部署医疗大语言模型中的幻觉与作用层面滥用

Sunday Oyinlola Ogundoyin, Muhammad Ikram, Rahat Masood

机构 * The University of New South Wales, Sydney, Australia(新南威尔士大学,悉尼,澳大利亚)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);分类 cs.CL、cs.CY

AI总结 本文研究了网络部署的医疗大语言模型中的幻觉和作用层面滥用问题,通过评估6233个MedGPT和10个开源LLM,发现25-30%的MedGPT事实准确性较低,33.6-54.3%违反操作阈值,57.06%的Action-enabled模型缺乏充分的隐私披露,揭示了系统性漏洞,强调了多指标评估和更强的安全保障的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26566 2026-04-22 cs.LG cs.AI 73%

Multiclass Local Calibration with the Jensen-Shannon Distance

多类局部校准与 Jensen-Shannon 距离

Cesare Barbera, Lorenzo Perini, Giovanni De Toni, Andrea Passerini, Andrea Pugnana

机构 * University of Trento(特伦托大学) University of Pisa(比萨大学) Meta Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出多类局部校准概念,通过 Jensen-Shannon 距离改进神经网络的局部校准能力,解决现有方法在稀疏区域的校准偏差问题。

Comments Accepted at AISTATS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07382 2026-04-14 cs.LG cs.AI 73%

Latent Structure of Affective Representations in Large Language Models

大型语言模型中情感表征的潜在结构

Benjamin J. Choi, Melanie Weber

机构 * Harvard University(哈佛大学)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 研究大型语言模型中情感表征的潜在几何结构,发现其与心理学中的valence-arousal模型一致,并展示了其在不确定性量化中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16889 2026-03-19 cs.CL cs.AI cs.SD eess.AS 73%

Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment

基于评分标准的语音大模型多维度多评分者二语阅读语音评估

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于评分标准的语音大模型微调方法,通过多维度评分标准和不确定性校准,提升二语阅读语音评估的可靠性与可解释性。

Comments Accepted to LREC 2026. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants, which is financed by the Dutch Research Council (NWO)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15674 2026-03-19 cs.AI cs.IT cs.LG math.IT stat.ML 73%

Theoretical Foundations of Latent Posterior Factors: Formal Guarantees for Multi-Evidence Reasoning

潜在后验因子的理论基础:多证据推理的正式保证

Aliyu Agboola Alege

机构 * Epalea

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种理论框架,通过变分自编码器将多异质证据转换为高斯潜在后验,并利用Sum-Product网络或神经聚合器进行聚合,提供多证据推理的正式保证。

Comments 30 pages, 8 figures, 10 tables. Theoretical characterization of the Latent Posterior Factors (LPF) framework for multi-evidence probabilistic reasoning, with formal guarantees and empirical validation

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02798 2026-03-04 cs.AI cs.CL 73%

Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification

基于指南的证据积累用于高风险代理验证

Yichi Zhang, Nabeel Seedat, Yinpeng Dong, Peng Cui, Jun Zhu, Mihaela van de Schaar

机构 * Tsinghua University(清华大学) University of Cambridge(剑桥大学) Thomson Reuters Foundational Research(汤姆森路透基础研究)

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 GLEAN通过基于指南的证据积累框架,提升高风险代理决策的验证可靠性,实验显示其在AUROC和Brier分数减少方面均优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05184 2026-02-06 hep-th cond-mat.dis-nn cs.AI cs.LG 73%

Towards Worst-Case Guarantees with Scale-Aware Interpretability

面向尺度感知可解释性的最坏情况保证

Lauren Greenspan, David Berman, Aryeh Brill, Ro Jefferson, Artemy Kolchinsky, Jennifer Lin, Andrew Mack, Anindita Maiti, Fernando E. Rosas, Alexander Stapleton, Lucas Teixeira, Dmitry Vaintrob

机构 * Principles of Intelligence, USA(智能原理研究所,美国) Centre for Theoretical Physics, Queen Mary University of London(理论物理中心,伦敦女王大学) Universitat Pompeu Fabra, Barcelona, Spain(庞培法布拉大学,巴塞罗那,西班牙) Perimeter Institute for Theoretical Physics, Waterloo ON, Canada(皮尔姆研究所,滑铁卢,加拿大) Department of Informatics, University of Sussex(信息学院, Sussex 大学) Department of Brain Sciences, Imperial College London(脑科学系,伦敦帝国学院) Centre for Eudaimonia and human flourishing, University of Oxford(幸福与人类繁荣中心,牛津大学)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过物理重整化框架开发具有鲁棒性和忠实性的尺度感知可解释性工具,以提升人工智能的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01769 2026-02-04 cs.LG cs.AI 73%

IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination

IRIS: 隐式奖励引导的内部筛选以缓解多模态幻觉

Yuanshuai Li, Yuping Yan, Jirui Han, Fei Ming, Lingjuan Lv, Yaochu Jin

机构 * Department of Artificial Intelligence, Westlake University, Hangzhou, China(人工智能系,西湖大学,杭州,中国) Sony Research, Sony(索尼研究,索尼)

专题命中 幻觉与事实性 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG

AI总结 IRIS通过隐式奖励引导内部筛选,有效缓解多模态大语言模型的幻觉问题,无需外部反馈且性能优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12471 2026-01-23 cs.CL cs.AI 73%

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12804 2026-01-21 cs.AI cs.LG 73%

SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability

SL-CBM: 通过语义局部性增强概念瓶颈模型以提升可解释性

Hanwei Zhang, Luo Cheng, Rui Wen, Yang Zhang, Lijun Zhang, Holger Hermanns

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 SL-CBM通过引入语义局部性增强概念瓶颈模型,提升可解释性和干预效果,同时保持分类准确度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06460 2026-01-13 cs.CV cs.AI cs.CL 73%

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

语气至关重要:语言语气对VLMs幻觉影响的研究

Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li

机构 * Department of Computer Science and Technology, Kean University(计算机科学与技术系,凯恩大学) Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了提示语气对VLMs幻觉的影响,通过Ghost-100数据集发现幻觉率与提示强度非线性相关,揭示模型在处理结构性强制时的局限性。

Comments 10 pages, 6 figures, WACV Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12012 2025-12-17 cs.CV cs.AI cs.CL cs.RO 73%

Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus

语义驱动:通过开放词汇锚定和神经符号视觉语言共识民主化长尾数据整理

Antonio Guillen-Perez

机构 * Independent Researcher(独立研究者)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 Semantic-Drive通过开放词汇锚定和神经符号视觉语言共识,提升自动驾驶中长尾数据整理的效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11028 2025-12-15 cs.CL cs.AI 73%

Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models

注意置信差距:大型语言模型中的过度自信、校准与干扰效应

Prateek Chhikara

机构 * University of Southern California(美国南加州大学)

专题命中 幻觉与事实性 :RLHF(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨了大型语言模型中的过度自信问题,通过引入干扰项显著改善校准,提出针对性的改进策略以提升模型可靠性。

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06938 2025-11-24 cs.LG cs.AI 73%

From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers

从噪声到叙述:追踪变换器中幻觉的起源

Praneet Suresh, Jack Stanley, Sonia Joseph, Luca Scimeca, Danilo Bzdok

机构 * Mila - Quebec AI Institute(魁北克人工智能研究所)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 研究通过分析变换器模型在输入不确定性下的行为,揭示幻觉的起源及预测方法,为AI安全和风险评估提供依据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00126 2025-11-04 cs.LG cs.AI 73%

Dynamic Model Selection for Trajectory Prediction via Pairwise Ranking and Meta-Features

Lu Bowen

机构 * Lu Bowen(路 Bowen)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09279 2025-10-14 cs.CV cs.AI cs.CL 73%

Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models

Anita Kriz, Elizabeth Laura Janes, Xing Shen, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克AI研究所)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to ICCV 2025 Workshop CVAMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17225 2025-10-07 cs.CL cs.AI 73%

SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation

Xiaqiang Tang, Yi Wang, Keyu Hu, Rui Xu, Chuang Li, Weigao Sun, Jian Li, Sihong Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Shanghai AI Lab(上海人工智能实验室) Tencent Hunyuan(腾讯文英)

专题命中 幻觉与事实性 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11625 2025-10-01 cs.LG cs.AI cs.CR 73%

Inducing Uncertainty on Open-Weight Models for Test-Time Privacy in Image Recognition

Muhammad H. Ashiq, Peter Triantafillou, Hung Yun Tseng, Grigoris G. Chrysos

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Warwick(沃里克大学)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17407 2025-09-23 cs.AI cs.CL 73%

Post-hoc Reward Calibration: A Case Study on Length Bias

Zeyu Huang, Zihan Qiu, Zili Wang, Edoardo M. Ponti, Ivan Titov

机构 * University of Edinburgh(爱丁堡大学) Alibaba Group(阿里巴巴集团) INF Technology(INF技术) University of Amsterdam(阿姆斯特丹大学)

专题命中 幻觉与事实性 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14660 2025-07-25 cs.AI cs.CL 73%

When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems

Qibing Ren, Sitao Xie, Longxuan Wei, Zhenfei Yin, Junchi Yan, Lizhuang Ma, Jing Shao

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Code is available at https://github.com/renqibing/MultiAgent4Collusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11574 2025-07-17 cs.LG cs.AI stat.ML 73%

Distribution-Free Uncertainty-Aware Virtual Sensing via Conformalized Neural Operators

Kazuma Kobayashi, Shailesh Garg, Farid Ahmed, Souvik Chakraborty, Syed Bahauddin Alam

机构 * The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格雷格尔工程学院) National Center for Supercomputing Applications(国家超级计算应用中心) Department of Applied Mechanics, Indian Institute of Technology Delhi(印度理工学院德里应用力学系) Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05246 2025-07-08 cs.AI cs.CL 73%

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors

Scott Emmons, Erik Jenner, David K. Elson, Rif A. Saurous, Senthooran Rajamanoharan, Heng Chen, Irhum Shafkat, Rohin Shah

机构 * Google(谷歌)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏