arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2410.02768 2025-05-07 cs.CV cs.AI 79%

Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Jin Chen, Kaijing Ma, Haojian Huang, Han Fang, Hao Sun, Mehdi Hosseinzadeh, Zhe Liu

机构 * School of Computer Science, Duy Tan University(计算机科学学院,杜益坦大学) School of Computer Sciences, Universiti Sains Malaysia(计算机科学学院,马来亚大学)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09445 2025-04-02 cs.CV cs.AI 79%

Astrea: A MOE-based Visual Understanding Model with Progressive Alignment

Xiaoda Yang, JunYu Lu, Hongshun Qiu, Sijing Li, Hao Li, Shengpeng Ji, Xudong Tang, Jiayang Xu, Jiaqi Duan, Ziyue Jiang, Cong Lin, Sihang Cai, Zejian Xie, Zhuoyang Song, Songxin Zhang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01289 2024-11-05 cs.LG 79%

Uncertainty measurement for complex event prediction in safety-critical systems

Maria J. P. Peixoto, Akramul Azim

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15686 2024-10-22 cs.MA cs.AI 79%

NetSafe: Exploring the Topological Safety of Multi-agent Networks

Miao Yu, Shilong Wang, Guibin Zhang, Junyuan Mao, Chenlong Yin, Qijiong Liu, Qingsong Wen, Kun Wang, Yang Wang

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05938 2024-10-10 cs.CV cs.AI 79%

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

Yifei Xing, Xiangyuan Lan, Ruiping Wang, Dongmei Jiang, Wenjun Huang, Qingfang Zheng, Yaowei Wang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13692 2024-10-07 cs.CL 79%

Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation

Di Wu, Jia-Chen Gu, Fan Yin, Nanyun Peng, Kai-Wei Chang

专题命中 幻觉与事实性 :trustworthy(title);alignment(abstract);分类 cs.CL

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11007 2024-09-18 cs.CL cs.CV 79%

CAST: Cross-modal Alignment Similarity Test for Vision Language Models

Gautier Dagan, Olga Loginova, Anil Batra

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01151 2024-09-04 cs.CV cs.LG 79%

Understanding Multimodal Hallucination with Parameter-Free Representation Alignment

Yueqian Wang, Jianxin Liang, Yuxuan Wang, Huishuai Zhang, Dongyan Zhao

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02449 2024-07-02 cs.LG stat.AP 79%

Bayesian Safety Validation for Failure Probability Estimation of Black-Box Systems

Robert J. Moss, Mykel J. Kochenderfer, Maxime Gariel, Arthur Dubois

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.LG

Journal ref AIAA Journal of Aerospace Information Systems (JAIS) 21.7 (2024): 533-546

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04854 2024-06-10 cs.CL 79%

Uncertainty Aware Learning for Language Model Alignment

Yikun Wang, Rui Zheng, Liang Ding, Qi Zhang, Dahua Lin, Dacheng Tao

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12646 2024-01-24 cs.MA cs.AI cs.GT 79%

Emergent Cooperation under Uncertain Incentive Alignment

Nicole Orzan, Erman Acar, Davide Grossi, Roxana Rădulescu

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13240 2023-11-23 cs.CL 79%

On the Calibration of Large Language Models and Alignment

Chiwei Zhu, Benfeng Xu, Quan Wang, Yongdong Zhang, Zhendong Mao

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments to be published in findings of EMNLP-2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06697 2023-11-14 cs.CL 79%

Trusted Source Alignment in Large Language Models

Vasilisa Bashlovkina, Zhaobin Kuang, Riley Matthews, Edward Clifford, Yennie Jun, William W. Cohen, Simon Baumgartner

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16351 2023-08-22 eess.SY cs.AI cs.SY 79%

Distributionally Robust Safety Filter for Learning-Based Control in Active Distribution Systems

Hoang Tien Nguyen, Dae-Hyun Choi

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04309 2023-03-31 cs.AI 79%

Alignment-based conformance checking over probabilistic events

Jiawei Zheng, Petros Papapanagiotou, Jacques D. Fleuriot

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

Comments Extended version

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11401 2022-12-05 cs.CL cs.CV cs.MM 79%

Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations

Qian Yang, Yunxin Li, Baotian Hu, Lin Ma, Yuxing Ding, Min Zhang

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL

Comments 11 pages (including Supplementary Materials); Accepted to ACM MM 2022

Journal ref ACM International Conference on Multimedia. 2022. 3587-3597

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03053 2021-02-08 cs.AI 79%

Risk-Constrained Interactive Safety under Behavior Uncertainty for Autonomous Driving

Julian Bernhard, Alois Knoll

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06622 2021-01-19 cs.LG stat.ML 79%

Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings

Jesse Zhang, Brian Cheung, Chelsea Finn, Sergey Levine, Dinesh Jayaraman

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.LG

Comments 15 pages, 8 figures, ICML 2020. Website with code: https://sites.google.com/berkeley.edu/carl

Journal ref Proceedings of the 37th International Conference on Machine Learning, PMLR 119:11055-11065, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.07907 2017-09-28 cs.AI 79%

Mutual Alignment Transfer Learning

Markus Wulfmeier, Ingmar Posner, Pieter Abbeel

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16591 2026-07-21 cs.LG cs.AI 新提交 79%

Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL

从世界反馈中学习:为何模型不确定性在基于模型的强化学习中无法作为风险信号

Zhaohui Wang

专题命中 幻觉与事实性 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 研究探讨 RLxF 中学习信号应源于世界反馈,在安全模型控制中实例化并提炼原则。通过实验表明基于动力学的不确定性惩罚会增加碰撞率,用世界反馈信号可降低碰撞率,提取原则并指出其适用于多种相关方法。

Comments Accepted at the ICML 2026 Workshop on Reinforcement Learning from X Feedback (RLxF). 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30381 2026-06-01 cs.LG cs.AI 79%

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

当LLM学会一致错误:合成欺骗的线性表示的多模型研究

Vahideh Zolfaghari

机构 * Algoverse AI Research Medical Sciences Education Research Center, Mashhad University of Medical Sciences(马什哈德大学医学科学教育研究中心) Student Research Committee, Department of Health Information Technology and Management, Medical Informatics, School of Allied Medical Sciences, Shahid Beheshti University of Medical Sciences(谢赫·贝赫什提大学医学科学学院学生研究委员会,健康信息科技与管理系,医学信息学)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 通过LoRA微调五个Transformer模型的诚实与欺骗变体,使用线性探针检测合成欺骗,发现早期层即可达到近完美AUC,支持线性表示假说,并揭示两种表示机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03871 2025-09-05 cs.CL cs.AI cs.CR 79%

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Yanbo Wang, Yongcan Yu, Jian Liang, Ran He

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(神经语言处理与机器智能中心,中国科学院自动化研究所)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 38 pages. This survey considers papers published up to June 30, 2025. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18129 2025-06-24 cs.CL cs.AI 79%

$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models

Bugra Kilictas, Faruk Alpay

机构 * Bahcesehir University(巴切希尔大学)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13746 2025-06-17 cs.CR cs.AI cs.LG 79%

Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability

Shova Kuikel, Aritran Piplai, Palvi Aggarwal

机构 * University of Texas at El Paso(德克萨斯理工大学)

专题命中 幻觉与事实性 :alignment(abstract);DPO(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07141 2026-08-10 cs.CV 新提交 78%

Human-AI Perceptual Alignment by Playing Hues and Cues

通过玩Hues and Cues游戏实现人类与AI的感知对齐

Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver, Pablo Hernández-Cámara, Valero Laparra, Jesús Malo

机构 * Universitat de València(瓦伦西亚大学) Image Processing Lab(图像处理实验室)

专题命中 幻觉与事实性 :alignment(title,abstract)

AI总结 本研究提出基于Hues and Cues游戏的评估框架,对比162个CVLMs与人类的颜色感知对齐,发现其在抽象领域偏离人类基线,筛选的预训练数据集可缓解错位。

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03460 2026-06-03 cs.CV 78%

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

从3D感知到安全推理:基于图的实时地下矿井监控框架

Pasindu Ranasinghe, Simit Raval, Dibyayan Patra, Bikram Banerjee, Ismet Canbulat

专题命中 幻觉与事实性 :safety(title,abstract)

AI总结 提出一个结合3D语义感知、不确定性异常检测、规则检查、设备端LLM推理和GraphRAG记忆分析的连续监控框架,通过场景图和时序图实现结构化安全推理,在115个危险场景中达到93%的覆盖率和92.7%的感知精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05249 2026-06-03 cs.IR 78%

TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation

TriAlignGR:面向生成式推荐的多模态深度兴趣挖掘与三角多任务对齐

Yangchen Zeng, Hao Peng, Rongfeng Guo, Zhenyu Yu, Zhiyuan Hu, Jinze Wang

专题命中 幻觉与事实性 :alignment(title,abstract)

AI总结 提出TriAlignGR统一多任务多模态框架,通过两阶段语义传播和三角多任务对齐,解决语义ID中的内容退化与语义不透明问题,实现生成式推荐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18838 2026-06-02 cs.LG cs.AI cs.CL 78%

Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

说谎只是一个阶段:语言模型扩展中的隐藏对齐转变

Adil Amin

机构 * ZEHEN Labs(ZEHEN实验室)

专题命中 幻觉与事实性 :alignment(title);分类 cs.CL、cs.AI、cs.LG

AI总结 通过分析63个基础模型,发现语言模型在特定规模阈值下,推理能力与真实性从反相关转变为正相关,并揭示了输出投影瓶颈和零竞争注意力头等内部机制。

Comments 15 pages, 8 figures, 2 tables. Companion paper: "The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next." ( https://doi.org/10.48550/arXiv.2605.18840). Code: https://github.com/adilamin89/cape-scaling. Dashboard: https://zehenlabs.com/cape/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18442 2026-05-19 cs.RO 78%

SG-CADVLM: A Context-Aware Decoding Powered Vision Language Model for Safety-Critical Scenario Generation

SG-CADVLM: 一种基于上下文感知解码的视觉语言模型,用于安全关键场景生成

Hongyi Zhao, Shuo Wang, Qijie He, Ziyuan Pu

机构 * School of Transportation, Southeast University(东南大学交通学院)

专题命中 幻觉与事实性 :safety(title,abstract)

AI总结 本文提出SG-CADVLM,一种结合上下文感知解码的多模态输入处理框架,用于从事故报告中生成高保真的安全关键场景,通过减少视觉语言模型的幻觉并同时生成道路几何和车辆轨迹,提升了生成场景的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20903 2026-04-24 cs.CR 78%

Sensitivity Uncertainty Alignment in Large Language Models

大语言模型中的敏感性不确定性对齐

Prakul Sunil Hiremath, Harshit R. Hiremath

专题命中 幻觉与事实性 :alignment(title,abstract)

AI总结 本文提出SUA框架,用于分析大语言模型在对抗性和模糊输入下的失败原因,通过敏感性与不确定性对齐提升模型可靠性。

Comments 24 pages, 4 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏