arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2301.06009 2023-01-18 cs.CL cs.AI 62%

Rationalizing Predictions by Adversarial Information Calibration

Lei Sha, Oana-Maria Camburu, Thomas Lukasiewicz

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2012.08884

Journal ref Artificial Intelligence, Volume 315, February 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01425 2022-12-06 cs.CL cs.AI 62%

Unveiling the Black Box of PLMs with Semantic Anchors: Towards Interpretable Neural Semantic Parsing

Lunyiu Nie, Jiuding Sun, Yanlin Wang, Lun Du, Lei Hou, Juanzi Li, Shi Han, Dongmei Zhang, Jidong Zhai

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments AAAI 2023 Main Track Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03041 2022-11-08 cs.CL cs.AI 62%

Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence Estimates

Dongfang Li, Baotian Hu, Qingcai Chen

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.04260 2022-10-19 cs.LG cs.AI cs.CV 62%

Provably Robust Detection of Out-of-distribution Data (almost) for free

Alexander Meinke, Julian Bitterwolf, Matthias Hein

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04714 2022-10-17 cs.CL cs.LG stat.ML 62%

Uncertainty Quantification with Pre-trained Language Models: A Large-Scale Empirical Analysis

Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted by EMNLP 2022 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11724 2022-09-21 cs.LG cs.AI cs.SI 62%

Explainable Misinformation Detection Across Multiple Social Media Platforms

Gargi Joshi, Ananya Srivastava, Bhargav Yagnik, Mohammed Hasan, Zainuddin Saiyed, Lubna A Gabralla, Ajith Abraham, Rahee Walambe, Ketan Kotecha

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 28 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11097 2022-06-29 cs.RO cs.AI cs.LG 62%

UMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning

Christopher Diehl, Timo Sievernich, Martin Krüger, Frank Hoffmann, Torsten Bertram

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Best Paper Award @ Advances in Neural Information Processing Systems - Machine Learning for Autonomous Driving Workshop (NeurIPS 2021 ML4AD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09546 2022-06-10 cs.LG cs.CY 62%

Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions

Eike Petersen, Yannik Potdevin, Esfandiar Mohammadi, Stephan Zidowitz, Sabrina Breyer, Dirk Nowotka, Sandra Henn, Ludwig Pechmann, Martin Leucker, Philipp Rostalski, Christian Herzog

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CY、cs.LG

Journal ref IEEE Access, Vol. 10, pp. 58375-58418, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04837 2022-03-10 eess.AS cs.CL cs.CY 62%

'Beach' to 'Bitch': Inadvertent Unsafe Transcription of Kids' Content on YouTube

Krithika Ramesh, Ashiqur R. KhudaBukhsh, Sumeet Kumar

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.CY

Comments This paper got accepted at AAAI 2022, AI for Social Impact track

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02173 2022-03-04 cs.RO cs.AI cs.CV cs.LG cs.MA 62%

Multi-Agent Variational Occlusion Inference Using People as Sensors

Masha Itkina, Ye-Ji Mun, Katherine Driggs-Campbell, Mykel J. Kochenderfer

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 9 figures, International Conference on Robotics and Automation (ICRA) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08239 2022-02-11 cs.CL cs.AI 62%

LaMDA: Language Models for Dialog Applications

Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, Quoc Le

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.07586 2021-05-05 cs.CY cs.HC cs.LG 62%

Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty

Umang Bhatt, Javier Antorán, Yunfeng Zhang, Q. Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Gauthier Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, Lama Nachman, Rumi Chunara, Madhulika Srikumar, Adrian Weller, Alice Xiang

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society (AIES) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10715 2021-04-23 cs.LG cs.AI 62%

Uncertainty-Aware Boosted Ensembling in Multi-Modal Settings

Utkarsh Sarawgi, Rishab Khincha, Wazeer Zulfikar, Satrajit Ghosh, Pattie Maes

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at IJCNN 2021, to appear in IEEE proceedings. Equal contributions from US, RK and WZ

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03613 2021-04-09 cs.LG cs.AI stat.AP 62%

Uncertainty-aware Remaining Useful Life predictor

Luca Biggio, Alexander Wieland, Manuel Arias Chao, Iason Kastanis, Olga Fink

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.07650 2021-02-12 stat.ML cs.AI cs.LG 62%

Uncertainty Estimation in Autoregressive Structured Prediction

Andrey Malinin, Mark Gales

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06523 2020-12-14 cs.CV cs.AI cs.LG 62%

Dependency Decomposition and a Reject Option for Explainable Models

Jan Kronenberger, Anselm Haselhoff

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at CVPR 2019 Workshop "DThree19: Dependable Deep Detectors"

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.00496 2020-04-17 cs.LG cs.AI stat.ML 62%

Uncertainty-Based Out-of-Distribution Classification in Deep Reinforcement Learning

Andreas Sedlmeier, Thomas Gabor, Thomy Phan, Lenz Belzner, Claudia Linnhoff-Popien

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:1901.02219

Journal ref Proceedings of the 12th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART, 2020, ISBN 978-989-758-395-7, pages 522-529

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09005 2019-11-21 cs.AI cs.CY cs.SY eess.SY 62%

Hard Choices in Artificial Intelligence: Addressing Normative Uncertainty through Sociotechnical Commitments

Roel Dobbe, Thomas Krendl Gilbert, Yonatan Mintz

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY

Comments To be presented at the AI for Social Good workshop at NeurIPS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11538 2019-09-26 cs.RO cs.AI cs.LG cs.SY eess.SY stat.ML 62%

Automated Lane Change Decision Making using Deep Reinforcement Learning in Dynamic and Uncertain Highway Environment

Ali Alizadeh, Majid Moghadam, Yunus Bicer, Nazim Kemal Ure, Ugur Yavas, Can Kurtulus

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE Intelligent Transportation Systems Conference - ITSC 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.10134 2018-08-31 cs.RO cs.AI cs.LG 62%

Baidu Apollo Auto-Calibration System - An Industry-Level Data-Driven and Learning based Vehicle Longitude Dynamic Calibrating Algorithm

Fan Zhu, Lin Ma, Xin Xu, Dingfeng Guo, Xiao Cui, Qi Kong

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18985 2025-07-30 cs.CV cs.AI 61%

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Guanxi Shen

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :alignment(abstract,comments);分类 cs.AI

Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.03438 2019-03-11 cs.AI cs.RO 61%

Towards a Framework to Manage Perceptual Uncertainty for Safe Automated Driving

Krzysztof Czarnecki, Rick Salay

专题命中 幻觉与事实性 :safety(abstract,journal_ref);分类 cs.AI

Journal ref In: Gallina B., Skavhaug A., Schoitsch E., Bitsch F. (eds) Computer Safety, Reliability, and Security. SAFECOMP 2018. Lecture Notes in Computer Science, vol 11094. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13563 2026-08-17 cs.HC cs.AI cs.SE 新提交 57%

Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support

代理验证的大语言用户体验微模拟:一种用于早期决策支持的工件优先协议

Alexandre Cristovão Maiorano

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

AI总结 该研究提出代理验证的LLM驱动UX微模拟流程,通过代理语料库验证模拟结果,结合多指标对比基线、消融实验分析智能体策略,提供可复现的早期UX决策支持方案。

Comments 27 pages, 5 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23982 2026-08-14 cs.MA cs.AI 版本更新 57%

Moral Hazard in Multi-Agent Language Models

多智能体语言模型中的道德风险

Dane Malenfant

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

AI总结 研究多智能体语言模型中的道德风险,引入对话道德风险游戏,评估七个模型并分解行为,用多种优化机制更新,发现效果异质,强调应报告机制级行为而非仅团队成功。

Comments This revision substantially expands the empirical evaluation to eleven open-weight and three frontier models, adding matched query-cost, team-reward, group-size, and private-share incentive analyses. It also extends the weight-level and GEPA results, frozen-prompt information-structure interventions, statistical uncertainty analyses, and qualitative prompt/reasoning-trace studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13658 2026-08-14 cs.LG 版本更新 57%

Post-Hoc Uncertainty-Aware Explanations for Deployed Power Quality Disturbance Classifiers via Laplace Approximation

基于不确定性的贝叶斯解释框架用于电力质量问题分类

Yinsong Chen, Samson S. Yu, Kashem M. Muttaqi

机构 * School of Engineering, Deakin University(德肯大学工程学院) ARC Training Centre in Energy Technologies for Future Grids, School of Engineering, University of Wollongong(未来电网能源技术培训中心,沃尔灵宗大学工程学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

AI总结 本文提出一种贝叶斯解释框架,通过生成相关性归因分布来建模解释不确定性,提升电力质量问题分类器的透明度和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11631 2026-08-13 cs.AI 新提交 57%

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

CLAIM:基于不确定性度量的大语言模型开放域主动澄清领先方案

Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

AI总结 本研究提出基于不确定性度量的CLAIM框架,无需人工标注,结合监督学习与强化学习训练统一澄清决策模型,实现开放域人机交互中低成本鲁棒的主动澄清。

Comments 11 pages, 4 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11175 2026-08-13 cs.AI 版本更新 57%

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

自我进化临床系统之路:将医疗智能体从辅助扩展到自主

Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Chee Wei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin

机构 * Hunan University(湖南大学) ByteDance(字节跳动) Duke University(杜克大学) Westlake University(西湖大学) The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Macau(澳门大学) The Ohio State University(俄亥俄州立大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 研究探讨大语言模型等对医疗智能体的重塑,从临床部署出发,将其形式化为决策系统并给出自主性分类。沿统一框架扩展,强调临床环境扩展为关键方向,定位临床自我进化为前沿,还研究了多领域应用及挑战,提供医学成像系统路线图。

Comments Project page: https://github.com/zhcz328/Awesome-Medical-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10964 2026-08-12 cs.CV cs.AI 新提交 57%

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

CARE:用于可靠医学视觉问答的置信感知推理

Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu

机构 * Ant Group(蚂蚁集团) University of Michigan(密歇根大学) City University of Hong Kong(香港城市大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

AI总结 该研究针对医学多模态大语言模型的置信度校准问题,提出CARE框架,通过双阶段流程优化准确率与校准度,在三个医学VQA基准上取得最优性能,为临床决策支持提供可信基础。

Comments Accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10903 2026-08-12 cs.CV cs.LG 新提交 57%

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

VIDS-Seg:面向儿科心脏超声分割的可靠不确定性量化

Paul Fischer, Ece Ozkan

机构 * University of Basel(巴塞尔大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

AI总结 该研究提出VIDS-Seg方法,在成人超声数据集训练后零样本应用于儿科心脏超声分割,可可靠检测模型在儿科亚群的隐性失效,兼具分割精度与不确定性匹配优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09011 2026-08-11 cs.LG 新提交 57%

Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning

视觉-语言表示学习中的动态分布感知不确定性跟踪

Ao Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang, Shufan Yang, Haoru Chen, Qing Gu

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室) ByteDance Inc(字节跳动公司)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

AI总结 针对视觉-语言模型事后不确定性量化方法忽略测试分布动态性的问题,提出DDA-UQ框架,通过高斯混合模型建模嵌入空间,实现动态不确定性估计,性能优于现有最优方法。

详情

展开后加载摘要…

URL PDF HTML 收藏