arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3266 篇

2503.10566 2026-02-10 cs.LG 70%

ASIDE: Architectural Separation of Instructions and Data in Language Models

ASIDE: 语言模型中指令与数据的架构分离

Egor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova, Soroush Tabesh, Sebastian Lapuschkin, Wojciech Samek, Christoph H. Lampert

机构 * Institute of Science and Technology Austria (ISTA)(奥地利科学与技术研究所) Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Tübingen AI Center(图宾根人工智能中心) Centre of eXplainable Artificial Intelligence(可解释人工智能中心) Technische Universität Berlin(柏林技术大学) Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所)

专题命中 安全训练 :safety(abstract);prompt injection(abstract);分类 cs.LG

AI总结 ASIDE通过在令牌嵌入层面实现指令与数据的分离,提升了语言模型的安全性和性能

Comments ICLR 2026 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17600 2026-02-04 cs.CY 70%

STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems

基于STAMP/STPA的AI系统失控因素特征化

Steve Barrett, Anna Bruvere, Sean P. Fillingham, Catherine Rhodes, Stefano Vergani

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 本文基于STAMP/STPA方法,提出一个结构化框架来特征化AI系统失控的因果因素,以提升AI系统安全运行的保障能力。

Comments This new version only corrects some typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17892 2026-01-27 cs.CY cs.AI cs.CL cs.LG 70%

Artificial Intelligence and Intellectual Property Rights: Comparative Transnational Policy Analysis

人工智能与知识产权权利:比较跨国政策分析

Sahibpreet Singh, Manjit Singh

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究分析了人工智能与知识产权法的结合,揭示印度法律在适应AI生成内容方面的不足,并提出协调法律分类以促进公平创新。

Comments Published in Journal of University Institute of Legal Studies, Vol. 19, Issue 1, pp. 182-208, 2025

Journal ref Journal of University Institute of Legal Studies 19(1), 182-208 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17168 2026-01-27 cs.AI cs.MA 70%

Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability

解释代理系统:超越模型解释到系统级问责

Judy Zhu, Dhari Gandhi, Himanshu Joshi, Ahmad Rezaie Mianroodi, Sedef Akinli Kocak, Dhanesh Ramachandran

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) University of Texas, Austin(德克萨斯大学奥斯汀分校) Dalhousie University(达尔豪斯大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文探讨了代理系统中可解释性技术的必要性,提出需设计专门方法以确保系统在目标形成、环境交互和结果评估等阶段的可追溯性和问责性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23473 2026-01-21 cs.AI 70%

EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions

EVOREFUSE: 进化提示优化用于评估和缓解大语言模型对伪恶意指令的过度拒绝

Xiaorui Wu, Fei Li, Xiaofeng Mao, Xin Zhang, Li Zheng, Yuxiang Peng, Chong Teng, Donghong Ji, Zhuang Li

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航天信息安全部门、教育部、武汉大学计算机科学与工程学院) Ant Group(蚂蚁集团) Ant International(蚂蚁国际) School of Computing Technologies, Royal Melbourne Institute of Technology(皇家墨尔本理工学院计算机技术学院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 EVOREFUSE通过进化算法生成多样化伪恶意指令,提升LLM对恶意指令的拒绝触发率并减少过度拒绝

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03994 2026-01-21 cs.LG 70%

Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs

无需训练的策略违规检测:通过激活空间白化在大语言模型中

Oren Rachmil, Avishag Shapira, Roy Betser, Itay Gershon, Omer Hofman, Asaf Shabtai, Yuval Elovici, Roman Vainshtein

机构 * Fujitsu Research of Europe(富士通欧洲研究机构) Ben-Gurion University of the Negev(贝内杰尔大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 本文提出一种无需训练的策略违规检测方法,通过激活空间白化技术在大语言模型中实现高效检测。

Comments Accepted to the AAAI 2026 Deployable AI (DAI) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20293 2026-01-06 cs.CL 70%

AprielGuard

AprielGuard:一个统一安全与对抗防御的8B参数模型

Jaykumar Kasundra, Anjaneya Praharaj, Sourabh Surana, Lakshmi Sirisha Chodisetty, Sourav Sharma, Abhigya Verma, Abhishek Bhardwaj, Debasish Kanhar, Aakash Bhagat, Khalil Slimi, Seganrasan Subramanian, Sathwik Tejaswi Madhusudhan, Ranga Prasad Chenna, Srinivas Sunkara

机构 * AprielGuard

专题命中 安全训练 :safety(abstract);prompt injection(abstract);分类 cs.CL

AI总结 AprielGuard是一个统一安全与对抗防御的8B参数模型,通过多任务训练在多步骤和推理场景中有效检测有害内容和对抗操纵。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06059 2025-12-30 cs.CY 70%

Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles

优先原则,原则其次:一种适应性解读有助于、诚实和无害原则

Yue Huang, Chujie Gao, Yujun Zhou, Kehan Guo, Xiangqi Wang, Or Cohen-Sasson, Max Lamparth, Xiangliang Zhang

专题命中 安全训练 :alignment(abstract);harmlessness(abstract);分类 cs.CY

AI总结 本文提出了一种适应性解读有助于、诚实和无害原则的方法,通过优先级顺序平衡不同维度,以提高AI对齐的伦理和操作有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14860 2025-12-18 cs.CR cs.AI 70%

Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks

代理AI的渗透测试:跨模型和框架的比较安全分析

Viet K. Nguyen, Mohammad I. Husain

专题命中 安全训练 :safety(abstract);prompt injection(abstract);分类 cs.AI

AI总结 本文通过测试五个代理AI模型和两个框架,揭示了代理AI在安全方面的显著差异,发现超过一半的恶意提示成功,提出了新的防御策略和部署建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01848 2025-12-02 cs.CL 70%

Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability

超越SFT:通过强化学习构建更安全且推理能力更强的大推理模型

Jinghan Jia, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 本文提出通过强化学习提升大推理模型的安全性和推理能力,克服了传统监督微调的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00706 2025-12-02 cs.CV cs.AI 70%

Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation

利用策略数据优化LVLMs以实现有效的幻觉缓解

Chengzhi Yu, Yifan Xu, Yifan Chen, Wenyi Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 安全训练 :alignment(abstract);DPO(abstract);分类 cs.AI

AI总结 本文提出利用策略数据优化LVLMs,通过幻觉分类器和动态加权DPO算法,显著降低幻觉率并提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09341 2025-11-26 cs.AI 70%

Access Controls Will Solve the Dual-Use Dilemma

访问控制将解决双重用途困境

Evžen Wybitul

机构 * ETH Zurich, Switzerland(苏黎世联邦理工学院)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出基于访问控制的概念框架,通过验证用户访问双重用途输出,以解决人工智能安全系统在双重用途请求中的决策困境。

Comments Accepted at ICML 2025 Workshop on Technical AI Governance (TAIG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16144 2025-10-21 cs.NI cs.AI cs.MA 70%

Agentic AI for Ultra-Modern Networks: Multi-Agent Framework for RAN Autonomy and Assurance

Sukhdeep Singh, Avinash Bhat, Shweta M, Subhash K Singh, Moonki Hong, Madhan Raj K, Kandeepan Sithamparanathan, Sunder A. Khowaja, Kapal Dev

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15017 2025-10-20 cs.CR cs.AI 70%

Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks

ChenYu Wu, Yi Wang, Yang Liao

机构 * Wuhan, China(中国武汉) The University of Tokyo(东京大学) Xi’an Jiaotong University(西安交通大学)

专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 6pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23441 2025-10-15 cs.CL 70%

Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Tsinghua University(清华大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08046 2025-10-10 cs.AI 70%

LinguaSim: Interactive Multi-Vehicle Testing Scenario Generation via Natural Language Instruction Based on Large Language Models

Qingyuan Shi, Qingwen Meng, Hao Cheng, Qing Xu, Jianqiang Wang

机构 * National Key R&D Program of China(中华人民共和国国家重点研发计划)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00051 2025-10-09 cs.CY 70%

Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models

Philip Quirke, Narmeen Oozeer, Chaithanya Bandi, Amir Abdullah, Jason Hoelscher-Obermaier, Jeff M. Phillips, Joshua Greaves, Clement Neo, Michael Lan, Fazl Barez, Shriyash Upadhyay

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CY

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04392 2025-10-07 cs.CL cs.AI cs.CY cs.LG 70%

Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards

Faisal Hamman, Chenyang Zhu, Anoop Kumar, Xujun Peng, Sanghamitra Dutta, Daben Liu, Alfy Samuel

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Capital One

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at NeurIPS 2025 Workshop on Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17601 2025-10-07 cs.CL 70%

Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs

Jiawei Kong, Hao Fang, Xiaochen Yang, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Ke Xu, Han Qiu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Department of Software Engineering, Harbin Institute of Technology(哈尔滨工业大学软件工程系) School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区计算机科学与技术学院) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19839 2025-09-25 cs.AI 70%

LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation

Huizhen Shu, Xuying Li, Zhuo Li

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 9-page NeurIPS 2025 preprint including 3 figures and 1 table, with additional appendix material. Prepared using the NeurIPS 2025 preprint template and compiled with pdfLaTeX. All references are included via the provided .bbl file. Figures are in PDF format. No external supplementary files. All necessary style files and images are included

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12621 2025-09-25 cs.CL cs.IR 70%

SAFE: Improving LLM Systems using Sentence-Level In-generation Attribution

João Eduardo Batista, Emil Vatai, Mohamed Wahib

机构 * RIKEN-CCS Kobe, Japan(日本神户RIKEN-CCS)

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.CL

Comments 30 pages (9 pages of content, 5 pages of references, 16 pages of supplementary material), 7 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02159 2025-09-24 cs.LG 70%

PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning

Dongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia, Tianle Zhang, Kaige Zhang

机构 * School of Computer Science, University of Beihang, Beijing, China(北京航空航天大学计算机学院) School of Computer Science, Chinese University of Hong Kong, Hongkong, China(香港中文大学计算机学院) JD Explore Academy, Beijing, China(京东探索研究院) North Automatic Control Institute, Taiyuan, China(太原北自动控制研究所)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15491 2025-09-22 cs.RO cs.AI cs.SY eess.SY 70%

Explainable AI-Enhanced Supervisory Control for Robust Multi-Agent Robotic Systems

Reza Pirayeshshirazinezhad, Nima Fathi

机构 * Visual Computing and Computational Media, Texas A&M University(视觉计算与计算媒体,德克萨斯A&M大学) Mechanical Engineering Department, Texas A&M University(机械工程系,德克萨斯A&M大学) Department of Marine Engineering Technology, Texas A&M University(海洋工程技术系,德克萨斯A&M大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04429 2025-09-22 cs.AI 70%

Activation Space Interventions Can Be Transferred Between Large Language Models

Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Michael Lan, Abir Harrasse, Amirali Abdullah

机构 * Nvidia Singapore University of Technology and Design(新加坡技术与设计大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 75 pages. Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15805 2025-09-17 cs.CL 70%

Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering

Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee

机构 * Chung-Ang University(Chung-Ang 大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08075 2025-09-11 cs.CL 70%

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

Flor Miriam Plaza-del-Arco, Paul Röttger, Nino Scherrer, Emanuele Borgonovo, Elmar Plischke, Dirk Hovy

机构 * LIACS, Leiden University(莱顿大学 LIACS) Bocconi University(博科尼大学) Independent Researcher(独立研究者) Boconni University(博科尼大学) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍兹中心)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06786 2025-09-09 cs.LG 70%

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World

Youbang Sun, Xiang Wang, Jie Fu, Chaochao Lu, Bowen Zhou

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02650 2025-09-04 cs.AI cs.GT q-bio.PE 70%

Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis

Henrique Correia da Fonseca, António Fernandes, Zhao Song, Theodor Cimpeanu, Nataliya Balabanova, Adeela Bashir, Paolo Bova, Alessio Buscemi, Alessandro Di Stefano, Manh Hong Duong, Elias Fernandez Domingos, Ndidi Bianca Ogbo, Simon T. Powers, Daniele Proverbio, Zia Ush Shamszaman, Fernando P. Santos, The Anh Han, Marcus Krellner

机构 * INESC-ID and Instituto Superior Técnico, Universidade de Lisboa(INESC-ID和里斯本大学理工学院) School of Computing, Engineering and Digital Technologies, Teesside University(计算、工程与数字技术学院,泰赛德大学) Biological and Environmental Sciences, University of Stirling(生物与环境科学,斯特林大学) School of Mathematics, University of Birmingham(数学学院,伯明翰大学) Luxembourg Institute of Science and Technology(卢森堡科学与技术研究院) Machine Learning Group, Université libre de Bruxelles(机器学习小组,布鲁塞尔自由大学) AI Lab, Vrije Universiteit Brussel(人工智能实验室,布鲁塞尔自由大学) Division of Computing Science and Mathematics, University of Stirling(计算科学与数学系,斯特林大学) Department of Industrial Engineering, University of Trento(工业工程系,特伦特大学) University of Amsterdam(阿姆斯特丹大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 10 Pages, 7 Figures, accepted in the ALIFE 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20151 2025-08-29 cs.AI 70%

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement

Yuanzhe Shen, Zisu Huang, Zhengkang Guo, Yide Liu, Guanxu Chen, Ruicheng Yin, Xiaoqing Zheng, Xuanjing Huang

专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15068 2025-08-22 cs.AI 70%

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner

Shuang Ao, Gopal Rumchurn

机构 * Shuang Ao Gopal Rumchurn

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏