arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2601.11886 2026-04-21 cs.CL 79%

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

忠实性与安全性:在反事实医学证据下评估LLM行为

Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Northeastern University(东北大学) MD Anderson Cancer Center(MD安德森癌症中心) Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL

AI总结 本文研究了在反事实医学证据下LLM的行为,构建了MedCounterFact数据集,发现模型在面对危险或不合理证据时仍提供自信回答,表明模型可能过度强调忠实性而忽视安全性。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13101 2026-04-16 cs.SE cs.AI 79%

Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety

在天空中建立信任:一种基于知识的LLM框架用于航空安全

Anirudh Iyengar, Alisa Tiselska, Dumindu Samaraweera, Hong Liu

机构 * Senior Member, IEEE(IEEE高级会员) IEEE

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI

AI总结 本文提出结合LLM和知识图谱的框架,提升航空安全分析的可信度,通过双阶段流程构建和验证安全知识,提高准确性和可追溯性。

Comments Initial version of a conference publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10189 2026-04-14 cs.CL 79%

FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness

FAITH:通过整合可信度和诚实度实现事实一致性

Xiaoning Dong, Chengyan Wu, Yajie Wen, Yu Chen, Yun Xue, Jing Zhang, Wei Xu, Bolei Ma

机构 * Tsinghua University, Institute for Interdisciplinary Information Sciences(清华大学,交叉信息研究院) Shanghai Qi Zhi Institute(上海期智研究院) South China Normal University, School of Electronic Science and Engineering(华南师范大学,电子科学与工程学院) Guangzhou Richstone Data Technologies Co., Ltd.(广州瑞驰数据技术有限公司) LMU Munich & Munich Center for Machine Learning(慕尼黑大学 & 慕尼黑机器学习中心)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

AI总结 FAITH通过整合可信度和诚实度信号,提升大语言模型的事实准确性。采用PPO算法优化奖励函数,并结合检索增强模块增强内外知识一致性。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08384 2026-04-10 eess.AS cs.AI 79%

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

TASU2:用于语音大语言模型对齐和低资源适应的可控CTC模拟

Jing Peng, Chenghao Wang, Yi Yang, Lirong Qian, Junjie Li, Yu Xi, Shuai Wang, Kai Yu

机构 * X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系X-LANCE实验室) MoE Key Lab of Artificial Intelligence(教育部人工智能重点实验室) Jiangsu Key Lab of Language Computing(江苏省语言计算重点实验室) AISpeech Ltd(思必驰科技股份有限公司) Nanjing University(南京大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 TASU2通过可控CTC模拟生成文本监督,提升语音大语言模型的对齐和低资源适应性能,优于TASU和文本细调等基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06066 2026-04-08 cs.CL 79%

From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection

从幻觉到结构雪球效应:约束解码在LLM反思中的对齐税

Hongxu Zhou

机构 * Saarland University(萨尔大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

AI总结 研究探讨了在开放式推理任务中,约束解码通过强制结构反思是否能避免错误传播,发现反而引发结构雪球效应,并揭示了约束解码的对齐税问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24258 2026-03-26 cs.CL 79%

Semantic Alignment across Ancient Egyptian Language Stages via Normalization-Aware Multitask Learning

通过感知-意识多任务学习实现古代埃及语言各阶段的语义对齐

He Huang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

AI总结 本文研究古代埃及语言四个历史阶段的词级语义对齐,通过联合训练共享字级分词器的紧凑编码器-解码器模型,结合掩码语言模型、翻译语言模型、序列到序列翻译和词性标注任务,提升跨阶段对齐效果。

Comments Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05839 2026-03-09 cs.MA cs.AI 79%

Evaluating LLM Alignment With Human Trust Models

评估大语言模型与人类信任模型的对齐性

Anushka Debnath, Stephen Cranefield, Bastin Tony Roy Savarimuthu, Emiliano Lorini

机构 * School of Computing, University of Otago, New Zealand(奥塔哥大学计算机学院,新西兰) IRIT, CNRS, Toulouse University, France(IRIT,法国国家科学研究中心,图卢兹大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 本文通过对比提示生成嵌入向量,评估了EleutherAI/gpt-j-6B对信任的内部表示,发现其最接近Castelfranchi社会认知模型。

Comments This paper will appear in the post-proceedings of ICAART 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04897 2026-03-09 cs.CL 79%

Can LLMs Capture Expert Uncertainty? A Comparative Analysis of Value Alignment in Ethnographic Qualitative Research

LLMs能否捕捉专家不确定性?对人类学质性研究中价值对齐的比较分析

Arina Kostina, Marios Dikaiakos, Alejandro Porcel, Tassos Stassopoulos

机构 * University of Cyprus(塞浦路斯大学) University of Cambridge(剑桥大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

AI总结 本文研究LLMs在人类学质性研究中捕捉专家不确定性的能力,发现Qwen在价值对齐上表现最佳,但需进一步探讨模型偏差问题。

Comments Accepted for a poster session at BIG.AI at MIT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08500 2026-02-17 cs.AI 79%

Large Language Models as Oracles for Ontology Alignment

大型语言模型作为本体对齐的预言机

Sviatoslav Lushnei, Dmytro Shumskyi, Severyn Shykula, Ernesto Jimenez-Ruiz, Artur d'Avila Garcez

机构 * City St George’s, University of London, UK(伦敦大学城市圣乔治学院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出利用大型语言模型作为预言机,通过验证高不确定性的对应关系子集,提升本体对齐任务的性能,在OAEI 2025中取得优异成绩。

Comments Paper accepted at the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), main conference. 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06746 2026-02-02 cs.CV cs.AI 79%

AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model Alignment

AlignGemini: 通过任务-模型对齐实现通用的AI生成图像检测

Ruoxin Chen, Jiahui Gao, Kaiqing Lin, Keyue Zhang, Yandan Zhao, Isabel Guan, Taiping Yao, Shouhong Ding

机构 * Tencent Youtu Lab East China University of Science Shenzhen University Hong Kong University of Science Department of XXX, University of YYY, Location, Country School of ZZZ, Institute of WWW, Location, Country

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 AlignGemini通过任务-模型对齐原则,结合语义和像素伪影检测,提升AI生成图像检测的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09277 2026-01-30 cs.CL 79%

NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment

NeuroFaith:通过内部表示对齐评估LLM自解释的忠实性

Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau, Sarath Chandar, Marie-Jeanne Lesot

机构 * Sorbonne Université, CNRS, LIP6(索邦大学、国家科学研究中心、LIP6实验室) Mila - Quebec AI Institute, Université de Montréal(魁北克人工智能研究院、蒙特利尔大学)

专题命中 幻觉与事实性 :alignment(title);trustworthy(abstract);分类 cs.CL

AI总结 NeuroFaith通过内部表示对齐评估LLM自解释的忠实性,提出了一种灵活的框架和线性探针以提高模型的解释可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10986 2026-01-19 cs.CL 79%

ZPD Detector: Data Selection via Capability-Difficulty Alignment for Large Language Models

ZPD检测器:基于能力-难度对齐的数据选择用于大语言模型

Bo Yang, Yunkui Chen, Lanfei Feng, Yu Zhang, Shijian Li

机构 * State Key Laboratory of Brain–Machine Intelligence(脑机智能国家重点实验室) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

AI总结 ZPD检测器通过能力-难度对齐动态选择高价值样本,提升大语言模型的数据利用效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10803 2026-01-19 cs.CR cs.AI cs.CV 79%

Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment

超越已知伪造:通过事后分布对齐实现AI生成图像的通用检测

Li Wang, Wenyu Chen, Xiangtao Meng, Zheng Li, Shanqing Guo

机构 * School of Cyber Science and Technology, Shandong University(网络安全科学与技术学院,山东大学) State Key Laboratory of Cryptography and Digital Economy Security, Shandong University(密码与数字经济安全国家重点实验室,山东大学) Shandong Key Laboratory of Artificial Intelligence Security, Shandong University(人工智能安全山东省重点实验室,山东大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出PDA框架,通过事后分布对齐实现对未知生成器AI图像的通用检测,实验显示其在16种生成模型上的检测准确率达96.69%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04277 2026-01-09 cs.LG 79%

Unlocking the Pre-Trained Model as a Dual-Alignment Calibrator for Post-Trained LLMs

解封预训练模型作为后训练LLMs的双对齐校准器

Beier Luo, Cheng Wang, Hongxin Wei, Sharon Li, Xuefeng Du

机构 * Department of Statistics and Data Science, Southern University of Science and Technology(统计与数据科学系,南方科技大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学) Department of Computer Sciences, University of Wisconsin-Madison(计算机科学系,威斯康星大学麦迪逊分校) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出Dual-Align方法,通过双对齐策略校正后训练LLMs的置信度漂移和过程漂移,提升校准性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17254 2026-01-07 cs.CV cs.AI 79%

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

干预所有路径:统一缓解跨对齐格式的大型视觉-语言模型幻觉

Jiaye Qian, Ge Zheng, Yuchen Zhu, Sibei Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ShanghaiTech University(上海科技大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种统一干预框架,通过分析不同路径间的相互作用,有效缓解跨对齐格式的LVLM幻觉问题。

Comments Accepted to NeurIPS 2025, Project Page: https://github.com/SooLab/AllPath

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17793 2025-11-25 cs.CV cs.LG 79%

Attention Guided Alignment in Efficient Vision-Language Models

注意力引导的高效视觉-语言模型

Shweta Mahajan, Hoang Le, Hyojin Park, Farzad Farhadzadeh, Munawar Hayat, Fatih Porikli

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出AGE-VLM,通过交错交叉注意力层和空间知识提取,减少高效视觉-语言模型中的幻觉问题。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21323 2025-10-27 cs.CV cs.LG 79%

VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set

Shufan Shen, Junshu Sun, Qingming Huang, Shuhui Wang

机构 * Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS(智能信息处理重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18454 2025-10-22 cs.CL 79%

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan, Dinesh Manocha

机构 * Centre for Responsible AI, IIT Madras(负责任人工智能中心,印度理工学院马德拉斯分校) University of Maryland, College Park(马里兰大学 College Park 分校) Princeton University(普林斯顿大学)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16257 2025-10-21 cs.CL 79%

Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback

Chu Fei Luo, Samuel Dahan, Xiaodan Zhu

机构 * Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute(电气与计算机工程系及创新实验室研究机构) Conflict Analytics Lab, Queen’s University(冲突分析实验室,女王大学) Cornell Law School(康奈尔法学院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments Findings of EMNLP 2025, 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23109 2025-09-30 cs.AI cs.CV 79%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23002 2025-09-30 stat.ML cs.LG 79%

Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty

Lingyou Pang, Lei Huang, Jianyu Lin, Tianyu Wang, Akira Horiguchi, Alexander Aue, Carey E. Priebe

机构 * Department of Statistics, University of California, Davis(加州大学戴维斯分校统计系) Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

Comments 26 pages including appendix; 3 figures and 5 tables. Under review for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10048 2025-09-15 cs.LG 79%

Uncertainty-Aware Tabular Prediction: Evaluating VBLL-Enhanced TabPFN in Safety-Critical Medical Data

Madhushan Ramalingam

机构 * Engineering University of Moratuwa Sri Lanka(穆塔瓦大学工程学院)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15442 2025-09-08 eess.AS cs.AI cs.SD 79%

Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets

Chenlin Liu, Minghui Fang, Patrick Zhang, Wei Zhou, Jie Gao, Jiqing Han

机构 * Harbin Institute of Technology, China(哈尔滨工业大学) Zhejiang University, China(浙江大学) Tsinghua University, Shenzhen, China(清华大学深圳研究院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15880 2025-07-23 cs.AI 79%

The Recursive Coherence Principle: A Formal Constraint on Scalable Intelligence, Alignment, and Reasoning Architecture

Andy E. Williams

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06868 2025-06-10 cs.AI 79%

Incorporating Failure of Machine Learning in Dynamic Probabilistic Safety Assurance

Razieh Arshadizadeh, Mahmoud Asgari, Zeinab Khosravi, Yiannis Papadopoulos, Koorosh Aslansefat

机构 * School of Computer Science, University of Hull(计算机科学学院,赫尔大学)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04141 2025-05-30 cs.CL 79%

Reducing Tool Hallucination via Reliability Alignment

Hongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang, Su Zhu, Da Ma, Ruisheng Cao, Lu Chen, Kai Yu

机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China.(上海交通大学计算机科学学院X-LANCE实验室) MoE Key Lab of Artificial Intelligence, Shanghai, China.(人工智能MoE重点实验室) Jiangsu Key Lab of Language Computing, Suzhou, China.(江苏语言计算重点实验室) Suzhou Laboratory, Suzhou, China.(苏州实验室) AISpeech Co., Ltd., Suzhou, China.(AISpeech公司)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19327 2025-05-27 cs.LG 79%

Paying Alignment Tax with Contrastive Learning

Buse Sibel Korkmaz, Rahul Nair, Elizabeth M. Daly, Antonio del Rio Chanona

机构 * Imperial College London(帝国理工学院伦敦分校) IBM Research Europe(IBM欧洲研究院)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12221 2025-05-27 cs.CL 79%

On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation

Xueru Wen, Jie Lou, Xinyu Lu, Ji Yuqiu, Xinyan Guan, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Debing Zhang, Le Sun

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院软件研究所信息处理实验室,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Xiaohongshu Inc(小红书公司)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments Accepted as ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11803 2025-05-26 cs.CL 79%

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models

Boyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, Kam-Fai Wong

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments Accepted in ACL2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16806 2025-05-23 cs.CL cs.IR 79%

Two-way Evidence self-Alignment based Dual-Gated Reasoning Enhancement

Kexin Zhang, Junlan Chen, Daifeng Li, Yuxuan Zhang, Yangyang Feng, Bowen Deng, Weixu Chen

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏