arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-21 至 2026-04-21 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 10 篇

2604.16541 2026-04-21 cs.CV 82%

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

BOOKAGENT:通过多智能体认知校准 orchestrate 安全意识的视觉叙述

Bo Gao, Chang Liu, Yuyang Miao, Siyuan Ma, Ser-Nam Lim

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Science and Technology of China(中国科学技术大学) Imperial College London(伦敦帝国理工学院) Nanyang Technological University(南洋理工大学) University of Central Florida(佛罗里达中央大学)

专题命中 幻觉与事实性 :safety(title,abstract);alignment(abstract)

AI总结 本文提出BOOKAGENT,一种安全意识的多智能体协作框架,用于高质量的视觉叙述生成,通过联合规划、脚本、插图和全局修复不一致,提升叙述连贯性和视觉一致性。

Comments 18 pages, Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11886 2026-04-21 cs.CL 79%

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

忠实性与安全性:在反事实医学证据下评估LLM行为

Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Northeastern University(东北大学) MD Anderson Cancer Center(MD安德森癌症中心) Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL

AI总结 本文研究了在反事实医学证据下LLM的行为,构建了MedCounterFact数据集,发现模型在面对危险或不合理证据时仍提供自信回答,表明模型可能过度强调忠实性而忽视安全性。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02904 2026-04-21 cs.CL cs.AI cs.LG 69%

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning

通过不确定性校准微调增强大语言模型的信任

Ranganath Krishnan, Piyush Khanna, Omesh Tickoo

机构 * Capital One, AI Labs(Capital One人工智能实验室) Wayve Technologies(Wayve技术公司) Intel Corporation(英特尔公司)

专题命中 幻觉与事实性 :trustworthy(abstract,comments);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种不确定性感知微调方法,用于提升大语言模型在自然语言生成任务中的不确定性估计能力,从而提高生成响应的可信度并减少幻觉现象。

Comments ICLR 2026 Trustworthy AI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17651 2026-04-21 cs.CV cs.RO 67%

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

以基础设施为中心的世界模型:弥合时间深度与空间广度以实现道路感知

Siyuan Meng, Chengbo Ai

机构 * Department of Civil and Environmental Engineering, University of Massachusetts Amherst(土木与环境工程系,马萨诸塞大学阿默斯特分校)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract)

AI总结 本文提出以基础设施为中心的世界模型,通过时空互补性提升道路感知能力,提出三阶段框架和双层架构,结合多模态数据引擎和开放源代码基础,推动基础设施理解交通。

Comments 18 pages, 7 tables, 1 figure, vision paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05073 2026-04-21 cs.AI 57%

Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities

大语言模型代理中的不确定性量化:基础、新兴挑战与机遇

Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Xuefeng Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学) Nanyang Technological University(南洋理工大学) University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

AI总结 本文探讨了大语言模型代理中不确定性量化的基础、挑战及未来方向,提出新的框架并分析了现实场景中的技术难题。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15690 2026-04-21 cs.AI stat.AP 57%

From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models

从被动度量到主动信号:不确定性量化在大语言模型中的演变角色

Jiaxin Zhang, Wendi Cui, Zhuohang Li, Lifu Huang, Bradley Malin, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究院) Intuit(Intuit公司) Vanderbilt University(范德比大学) University of California, Davis(加州大学戴维斯分校) Vanderbilt University Medical Center(范德比大学医学中心)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 本文探讨大语言模型中不确定性量化从被动诊断指标到主动控制信号的演变,分析其在高级推理、自主代理和强化学习中的应用及贡献。

Comments This paper has been accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16929 2026-04-21 cs.CL 57%

MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

MeasHalu:通过增强推理缓解大型语言模型中的科学测量幻觉

Ruijun Huang, Zhiqiao Kang, Yuxuan Zhu, Junxiong Li, Jiahao Zhao, Minghuan Tan, Feng Jiang, Min Yang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(深圳先进技术大学人工智能研究院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

AI总结 本文提出MeasHalu框架,通过增强推理和针对性优化缓解LLM的科学测量幻觉问题,改进测量提取的准确性。

Comments To appear in ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17852 2026-04-21 cs.SD 50%

LLM-Codec: Neural Audio Codec Meets Language Model Objectives

LLM-Codec:神经音频编解码器与语言模型目标的结合

Ho-Lam Chung, Yiming Chen, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所) NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(国立清华大学人工智能研究中心) ASUS Intelligent Cloud Services(ASUS智能云服务)

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 本文提出LLM-Codec,通过引入未来token预测和语义对齐,提升音频编解码器与语言模型的协同性能,实验表明在语音连贯性和语音Mel距离上均有显著提升。

Comments ACL2026 Finding

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04814 2026-04-21 eess.IV cs.CV 50%

Explaining Uncertainty in Multiple Sclerosis Cortical Lesion Segmentation Beyond Prediction Errors

解释多发性硬化皮质病变分割中不确定性的原因超越预测误差

Nataliia Molchanova, Pedro M. Gordaliza, Alessandro Cagol, Mario Ocampo--Pineda, Po--Jui Lu, Matthias Weigel, Xinjie Chen, Erin S. Beck, Haris Tsagkas, Daniel Reich, Anna Stölting, Pietro Maggi, Delphine Ribes, Adrien Depeursinge, Cristina Granziera, Henning Müller, Meritxell Bach Cuadra

机构 * Faculty of Biology and Medicine, University of Lausanne (UNIL)(日内瓦大学生物医学学院) Radiology Department, Lausanne University Hospital (CHUV)(拉索恩大学医院放射科) MedGIFT, Institute of Informatics, School of Management, HES--SO Valais--Wallis University of Applied Sciences and Arts Western Switzerland(应用科学与艺术西瓦利斯-瓦利斯大学信息学院) CIBM Center for Biomedical Imaging(生物医学成像中心) Department of Radiology and Medical Informatics, University of Geneva(日内瓦大学放射科与医学信息学系) Translational Imaging in Neurology (ThINK) Basel, Department of Medicine and Biomedical Engineering, University Hospital Basel and University of Basel(神经学转化成像(ThINK)巴塞尔,医学与生物医学工程系,巴塞尔大学医院和巴塞尔大学) Multiple Sclerosis Center, Department of Neurology, University Hospital Basel(多发性硬化中心,神经科,巴塞尔大学医院) Research Center for Clinical Neuroimmunology and Neuroscience Basel (RC2NB), University Hospital Basel and University of Basel(临床神经免疫学和神经科学巴塞尔研究中心(RC2NB),巴塞尔大学医院和巴塞尔大学) Department of Neurology, Icahn School of Medicine at Mount Sinai(伊坎医学院 Mount Sinai 神经科) Translational Neuroradiology Section, National Institute of Neurological Disorders and Stroke, National Institutes of Health(神经学转化放射学部门,国家神经疾病与中风研究所,国家卫生研究院) Neuroinflammation Imaging Lab (NIL), Université catholique de Louvain(神经炎症成像实验室(NIL),卢瓦纳大学) EPFL+ECAL Lab, École polytechnique fédérale de Lausanne (EPFL)(EPFL+ECAL 实验室,日内瓦联邦理工学院(EPFL))

专题命中 幻觉与事实性 :trustworthy(abstract)

AI总结 本文提出一个框架,用于分析多发性硬化皮质病变分割中病变尺度的预测不确定性,揭示实例级不确定性与病变大小、形状和皮质涉及程度的关系,通过专家反馈验证其临床相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16952 2026-04-21 cs.CV 50%

Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder

更少的协同,更好的表现:通过条件和降质掩码自编码器解决异质多模态图像联合预训练

Bowen Peng, Yongxiang Liu, Jie Zhou, Xiaodong Chen, Tianpeng Liu, Xiaogang Yu, Li Liu

机构 * College of Electronic Science and Technology, National University of Defense Technology (NUDT)(电子科学与技术学院,国防科技大学) Beijing Institute of Remote Sensing Information(遥感信息研究所)

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 本文提出CoDe-MAE,通过Optical-anchored Knowledge Distillation和Conditioned Contrastive Learning解决高分辨率多模态联合预训练中的异质性-分辨率悖论,有效防止表示退化并在多个下游任务中取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏