arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1730 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1730 篇

2605.11398 2026-05-13 cs.AI cs.CL 81%

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

AcuityBench:评估语言模型对医疗紧急情况识别和不确定性对齐

Robin Linzmayer, Georgianna Lin, Di Coneybeare, Jason Chu, Trudi Cloyd, Manish Garg, Miles Gordon, Elizabeth Hartofilis, Benjamin Hong, Ashraf Hussain, Eugene Y. Kim, Oluchi Iheagwara King, Ross McCormack, Erica Olsen, John K. Riggins, Mustafa N. Rasheed, Dana L. Sacco, Vinay Saggar, Osman R. Sayan, Amit Shembekar, Janice Shin-Kim, Wendy W. Sun, Bernard P. Chang, David Kessler, Noémie Elhadad

机构 * Department of Computer Science, Columbia University, New York, NY, USA(计算机科学系,哥伦比亚大学,纽约,纽约州,美国) Department of Biomedical Informatics, Columbia University, New York, NY, USA(生物医学信息学系,哥伦比亚大学,纽约,纽约州,美国) Department of Emergency Medicine, Columbia University Irving Medical Center, New York, NY, USA(急诊医学系,哥伦比亚大学伊文思医疗中心,纽约,纽约州,美国)

专题命中 幻觉与事实性 :alignment(title);safety(abstract);分类 cs.CL、cs.AI

AI总结 AcuityBench通过统一框架评估语言模型对医疗紧急程度的识别能力,包含914个案例,涵盖明确和模糊情况,揭示模型在不同任务格式下的表现差异及不确定性处理问题。

Comments 41 pages, 5 figures. Preprint under review for the Track on Evaluations and Datasets at NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01643 2026-05-12 cs.LG cs.AI 81%

AI Alignment via Incentives and Correction

通过激励与修正实现AI对齐

Rohit Agarwal, Joshua Lin, Mark Braverman, Elad Hazan

机构 * Princeton University(普林斯顿大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文从法律经济学威慑模型视角探讨AI对齐问题,提出通过激励机制和修正过程解决AI行为与对齐目标的平衡问题,构建双代理模型分析奖励设计对求解器和审计器行为的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26959 2026-05-01 cs.CY cs.AI cs.MA 81%

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

CareGuardAI:面向临床安全与幻觉抑制的上下文感知多智能体守卫机制

Elham Nasarian, Abhilash Neog, Kwok-Leung Tsui, Niyousha HosseiniChimeh

机构 * Grado Department of Industrial & Systems Engineering, Virginia Tech, Blacksburg, VA 24061, USA(弗吉尼亚理工大学格拉多工业与系统工程系,弗吉尼亚理工大学,布莱克斯堡,VA 24061,美国) Department of Computer Science, Virginia Tech, Blacksburg, VA 24061, USA(弗吉尼亚理工大学计算机科学系,弗吉尼亚理工大学,布莱克斯堡,VA 24061,美国) Dept of Industrial, Manufacturing, and Systems Engineering at University of Texas at Arlington, Arlington, TX 76019, USA(德克萨斯理工大学阿灵顿分校工业、制造与系统工程系,阿灵顿,TX 76019,美国)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 CareGuardAI通过引入临床安全风险评估和幻觉风险评估,结合多阶段管道和迭代优化,提升患者面对LLM的可靠性与安全性,优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07965 2026-04-10 cs.CV cs.AI cs.LG 81%

DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editing

DSCA: 动态子空间概念对齐用于终身视觉语言模型编辑

Gyanendra Das, Sai Satyam Jena

机构 * Zynix AI, FL, USA(Zynix AI(美国佛罗里达州))

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 DSCA通过动态子空间对齐技术,在视觉语言模型中实现精准非干扰编辑,提升终身学习的稳定性与知识保留能力。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23783 2026-03-27 cs.LG cs.AI math.OC math.PR stat.ML 81%

Probabilistic Geometric Alignment via Bayesian Latent Transport for Domain-Adaptive Foundation Models

通过贝叶斯潜在传输的概率几何对齐用于领域自适应基础模型

Aueaphum Aueawatthanaphisut, Kuepon Auewattanapisut

机构 * School of Information, Computer Communication Technology Sirindhorn International Institute of Technology, Thammasat University Pathumthani, Thailand 0009-0006-4313-7359 epartment of Architecture, Faculty of Architecture Khon Kaen University Khon Kaen, Thailand

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种概率潜在传输框架,通过在表示空间中将领域适应建模为随机几何对齐问题,解决领域自适应中的潜在分布不匹配、优化动态不稳定和不确定性传播校准问题。

Comments 11 pages, 8 Figures, 25 Equations, 5 Tables and 3 Theorems

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12002 2026-01-21 cs.AI cs.LG cs.SY eess.SY 81%

Kernel-Based Learning of Safety Barriers

基于核的方法的安全屏障学习

Oliver Schön, Zhengang Zhong, Sadegh Soudjani

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于核的方法,用于安全屏障学习,通过数据驱动方法和控制屏障证书,提升安全验证的鲁棒性和可扩展性。

Comments 44 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03666 2026-01-12 cs.CL cs.AI cs.CV 81%

e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings

e5-omni: 显式跨模态对齐用于多模态嵌入

Haonan Chen, Sicheng Gao, Radu Timofte, Tetsuya Sakai, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) University of Würzburg(乌尔姆大学) Waseda University(早稻田大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 e5-omni通过显式对齐方法改进多模态嵌入,解决相似性尺度不一致、负样本效果下降和跨模态统计不匹配问题。

Comments https://huggingface.co/Haon-Chen/e5-omni-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09148 2025-12-11 cs.CL cs.AI 81%

Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment

通过注意力模式和语义对齐检测图检索增强生成中的幻觉

Shanghao Li, Jinda Han, Yibo Wang, Yuanjie Zhu, Zihe Song, Langzhou He, Kenan Kamel A Alghythee, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过注意力模式和语义对齐检测GraphRAG中的幻觉,开发了轻量级检测器GGA,提升了系统可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00091 2025-11-18 cs.AI cs.CL 81%

Ensemble Debates with Local Large Language Models for AI Alignment

Ephraiem Sarabamoun

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments The manuscript is being withdrawn to incorporate additional revisions and improvements

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19476 2025-10-23 cs.LG cs.AI 81%

A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring

Julian Schulz

机构 * Meridian Research, Cambridge(梅迪安研究,剑桥)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04045 2025-10-07 cs.CL cs.LG 81%

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

机构 * Columbia University(哥伦比亚大学) Barnard College(巴纳德学院)

专题命中 幻觉与事实性 :alignment(title);safety(abstract);分类 cs.CL、cs.LG

Comments ACL EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07974 2025-09-15 cs.SE cs.AI cs.LG 81%

From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems

Yining Hong, Christopher S. Timperley, Christian Kästner

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication at the International Conference on AI Engineering (CAIN) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01881 2025-08-26 cs.AI cs.CL 81%

WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue

Yaoyao Qian, Jindan Huang, Yuanli Wang, Simon Yu, Kyrie Zhixuan Zhou, Jiayuan Mao, Mingfu Liang, Hanhan Zhou

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Project website: https://nanostorm.netlify.app/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19548 2025-07-29 cs.CY cs.AI 81%

Justifications for Democratizing AI Alignment and Their Prospects

André Steingrüber, Kevin Baum

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Center for European Research in Trusted Artificial Intelligence (CERTAIN)(欧洲可信人工智能研究中心)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments accepted for the LNCS on-site proceedings of the AISoLA 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15906 2025-07-23 cs.LG cs.AI 81%

Towards Reliable, Uncertainty-Aware Alignment

Debangshu Banerjee, Kintan Saha, Aditya Gopalan

机构 * Undergraduate Programme, Indian Institute of Science, India(印度科学研究所)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14021 2025-06-09 cs.CL cs.LG q-bio.QM 81%

HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment

Yongqiang Chen, Quanming Yao, Juzheng Zhang, James Cheng, Yatao Bian

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICML2025, 27 pages, 7 figures, 23 tables; project page: https://higraphllm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22803 2025-05-30 cs.LG cs.AI 81%

CLUE: Neural Networks Calibration via Learning Uncertainty-Error alignment

Pedro Mendes, Paolo Romano, David Garlan

机构 * Software and Societal Systems Department, Carnegie Mellon University(卡内基梅隆大学软件与社会系统部门) INESC-ID and Instituto Superior Técnico, Universidade de Lisboa(里斯本大学INESC-ID和理工学院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20487 2025-05-28 cs.CL cs.AI 81%

InFact: Informativeness Alignment for Improved LLM Factuality

Roi Cohen, Russa Biswas, Gerard de Melo

机构 * Hasso Plattner Institute University of Potsdam(霍普夫纳研究所波茨坦大学) Dept. of Computer Science Aalborg University(计算机科学系奥尔堡大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01205 2025-04-03 cs.HC cs.AI cs.CL 81%

Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery

Nicholas Clark, Hua Shen, Bill Howe, Tanushree Mitra

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15079 2025-02-24 cs.CV cs.AI cs.CL 81%

Can Hallucination Correction Improve Video-Language Alignment?

Lingjun Zhao, Mingyang Xie, Paola Cascante-Bonilla, Hal Daumé, Kwonjoon Lee

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14913 2025-02-24 cs.CL cs.AI cs.IR 81%

OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment

Xiangjin Xie, Guangwei Xu, Lingyan Zhao, Ruijie Guo

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00352 2025-02-11 cs.CL cs.LG 81%

Does Alignment Tuning Really Break LLMs' Internal Confidence?

Hongseok Oh, Wonseok Hwang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Presented at the BlackboxNLP Workshop at EMNLP 2024 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04690 2024-12-09 cs.CL cs.AI 81%

LLM-Align: Utilizing Large Language Models for Entity Alignment in Knowledge Graphs

Xuan Chen, Tong Lu, Zhichun Wang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00219 2024-10-23 cs.CL cs.AI 81%

Evaluating Human Alignment and Model Faithfulness of LLM Rationale

Mohsen Fayyaz, Fan Yin, Jiao Sun, Nanyun Peng

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08985 2024-10-22 cs.AI cs.CL 81%

Towards Trustworthy Knowledge Graph Reasoning: An Uncertainty Aware Perspective

Bo Ni, Yu Wang, Lu Cheng, Erik Blasch, Tyler Derr

专题命中 幻觉与事实性 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01691 2024-10-03 cs.CL cs.AI 81%

FactAlign: Long-form Factuality Alignment of Large Language Models

Chao-Wei Huang, Yun-Nung Chen

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04910 2024-09-23 cs.CL cs.AI 81%

Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering

Weihe Zhai, Arkaitz Zubiaga

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15562 2024-08-29 cs.CL cs.LG 81%

Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation

Lujun Gui, Bin Xiao, Lei Su, Weipeng Chen

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments The work was not submitted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12416 2024-06-28 cs.CL cs.AI 81%

Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models

Hongbang Yuan, Yubo Chen, Pengfei Cao, Zhuoran Jin, Kang Liu, Jun Zhao

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13669 2024-06-14 cs.CL cs.AI 81%

The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models

Shuo Zhang, Liangming Pan, Junzhou Zhao, William Yang Wang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2024, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏