arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1715 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1715 篇

2605.09902 2026-05-12 cs.CV 82%

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

通过渐进分辨率处理和自适应特征对齐对多模态大语言模型进行对抗攻击

Haobo Wang, Xiaorong Ma, Weiqi Luo, Xiaojun Jia, Jiwu Huang

机构 * Sun Yat-sen University(中山大学) Nanyang Technological University(南洋理工大学) Shenzhen MSU-BIT University(深圳MSU-BIT大学)

专题命中 越狱攻击 :alignment(title,abstract);safety(abstract)

AI总结 本文提出PRAF-Attack框架,通过多尺度语义指导和中间层局部对齐提升多模态大语言模型的对抗攻击转移性,实验表明其在多种黑盒MLLM上表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00699 2026-05-08 cs.CR 82%

STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack

STARE:分步时间对齐与红队引擎用于多模态毒性攻击

Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng, Hongying Zan, Cong Wang

专题命中 越狱攻击 :alignment(title,abstract);safety(abstract)

AI总结 STARE通过分步时间对齐和红队引擎,提升多模态毒性攻击的成功率,揭示优化诱导的相位对齐现象,为安全机制提供理论基础。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10294 2026-04-28 cs.CR 82%

Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models

推理劫持:大语言模型中推理对齐的脆弱性

Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tun

专题命中 越狱攻击 :alignment(title,abstract);safety(abstract)

AI总结 本文指出现有安全研究忽视了推理对齐的脆弱性,提出新的对抗性提示攻击方法'推理劫持',通过注入虚假决策标准影响模型推理逻辑,实验表明即使先进模型也易受此类攻击影响。

Comments accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12371 2026-04-16 cs.CV 82%

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

在像素之间阅读:将文本-图像嵌入对齐与字体攻击成功性联系起来

Ravikumar Balakrishnan, Sanket Mendapara, Ankit Garg

机构 * Cisco Systems(思科系统)

专题命中 越狱攻击 :alignment(title);safety(abstract);prompt injection(abstract)

AI总结 研究了视觉-语言模型中字体提示注入攻击的影响,发现字体大小、文本攻击有效性及嵌入距离对攻击成功率有显著影响,为防御策略提供实证指导。

Comments Accepted at ICLR 2026 Workshop on Agents in the Wild

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05853 2026-04-09 cs.CV 82%

Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models

在像素之间阅读:一种针对文本到图像模型的 inscription 性 jailbreak 攻击

Zonghao Ying, Haowen Dai, Lianyu Hu, Zonglei Jing, Quanchen Zou, Yaodong Yang, Aishan Liu, Xianglong Liu

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

AI总结 本文提出 Etch 框架,通过分解对抗提示为语义伪装、视觉空间锚定和字形编码三层,实现对文本到图像模型的 inscription 性 jailbreak 攻击,验证了现有安全机制的不足。

Comments Withdrawn for extensive revisions and inclusion of new experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03598 2026-04-07 cs.CR 82%

AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models

AttackEval: 一种针对大语言模型的提示注入攻击有效性系统性实证研究

Jackson Wang

专题命中 越狱攻击 :prompt injection(title,abstract);safety(abstract)

AI总结 AttackEval系统性研究提示注入攻击有效性,揭示了Obfuscation攻击在对抗意图感知防御时的成功率最高,以及复合攻击策略显著提升攻击成功率,为改进大语言模型安全系统提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09246 2026-03-11 cs.CR 82%

Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models

面向推理的编程:通过链式语义工具来突破大型视觉语言模型

Quanchen Zou, Moyang Chen, Zonghao Ying, Wenzhuo Xu, Yisong Xiao, Deyue Zhang, Dongdong Yang, Zhao Liu, Xiangzheng Zhang

专题命中 越狱攻击 :jailbreak(title);alignment(abstract);safety(abstract)

AI总结 本文提出面向推理的编程方法,通过链式语义工具突破大型视觉语言模型的安全对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19610 2026-02-26 cs.CV 82%

JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models

JailBound: 视觉语言模型内部安全边界的劫持

Jiaxin Song, Yixu Wang, Jie Li, Rui Yu, Yan Teng, Xingjun Ma, Yingchun Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Xuan Tong(宣通) NSFOCUS

专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract)

AI总结 JailBound通过在视觉语言模型的潜在空间中探索安全边界,提出了一种新的劫持框架,有效提升了白盒和黑盒攻击成功率,揭示了模型的安全风险。

Comments The Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22240 2026-02-02 cs.CR cs.AI cs.CL cs.LG 82%

A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy

针对提示注入和劫持攻击的LLM防御措施系统文献综述:扩展NIST分类

Pedro H. Barcha Correia, Ryan W. Achjian, Diego E. G. Caetano de Oliveira, Ygor Acacio Maria, Victor Takashi Hayashi, Marcos Lopes, Charles Christian Miers, Marcos A. Simplicio

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文系统回顾了针对提示注入和劫持攻击的LLM防御措施,扩展了NIST分类,并提供了全面的防御目录和指南。

Comments 27 pages, 14 figures, 11 tables, submitted to Elsevier Computer Science Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22169 2026-02-02 cs.CL cs.AI cs.CR cs.LG 82%

In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement

葡萄酒中的真理与漏洞:通过醉酒语言诱导检验LLM安全性

Anudeex Shetty, Aditya Joshi, Salil S. Kanhere

机构 * School of Computer Science and Engineering, UNSW Sydney(计算机科学与工程学院,新南威尔士大学悉尼分校) School of Computing and Information System, the University of Melbourne(计算与信息系统学院,墨尔本大学)

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过诱导LLM产生醉酒语言,研究其安全漏洞,发现其易受劫持攻击和隐私泄露,揭示了LLM安全性的潜在风险。

Comments WIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03356 2026-01-13 cs.CR 82%

From static to adaptive: immune memory-based jailbreak detection for large language models

从静态到适应性:基于免疫记忆的大型语言模型 jailbreak 检测

Jun Leng, Yu Liu, Litian Zhang, Ruihan Hu, Zhuting Fang, Xi Zhang

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

AI总结 基于免疫记忆的框架,通过动态记忆更新和主动免疫机制,提升大型语言模型对 jailbreak 攻击的检测准确率至 94%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01513 2025-12-04 cs.CR cs.CV 82%

SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism

SafePTR: 通过剪枝-恢复机制实现多模态大语言模型的令牌级 Jailbreak 防御

Beitao Chen, Xinyu Lyu, Lianli Gao, Jingkuan Song, Heng Tao Shen

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳研究院) Southwestern University of Finance and Economics(西南财经大学) Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心) Tongji University(同济大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

AI总结 SafePTR 提出一种无需训练的多模态大语言模型防御机制,通过剪枝有害令牌并恢复良性特征,有效提升安全性并保持效率。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14140 2025-11-19 cs.CR 82%

Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security

Hajun Kim, Hyunsik Na, Daeseon Choi

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22564 2025-11-18 cs.CL cs.AI cs.LG 82%

Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs

Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07835 2025-10-10 cs.LG cs.AI cs.CL cs.CR 82%

MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation

Weisen Jiang, Sinno Jialin Pan

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted By NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21400 2025-09-29 cs.CR 82%

SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models

Xiyu Zeng, Siyuan Liang, Liming Lu, Haotian Zhu, Enguang Liu, Jisheng Dang, Yongbin Zhou, Shuchao Pang

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16870 2025-09-23 cs.SE cs.CR 82%

DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems

Rui Yang, Michael Fu, Chakkrit Tantithamthavorn, Chetan Arora, Gunel Gulmammadova, Joey Chua

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11630 2025-09-23 cs.CR cs.AI cs.CL cs.CY 82%

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine

机构 * Berkeley(伯克利) Mila – Quebec AI Institute(魁北克人工智能研究所) Montreal, Quebec, Canada(魁北克省蒙特利尔市) McGill University(麦吉尔大学) Georgia Tech(佐治亚理工学院)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15734 2025-06-23 cs.AI cs.CL cs.CR cs.CV cs.LG 82%

The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

Peiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun, Qin Xia, Zijiang Yang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院)

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00473 2025-06-19 cs.CV 82%

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, Tianxing He

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Shanghai Qi Zhi Institute(上海启智研究所) University of Chicago(芝加哥大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06679 2025-06-18 cs.CV 82%

T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks

Jiayang Liu, Siyuan Liang, Shiqian Zhao, Rongcheng Tu, Wenbo Zhou, Aishan Liu, Dacheng Tao, Siew Kei Lam

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16750 2025-06-13 cs.CR 82%

Guardians of the Agentic System: Preventing Many Shots Jailbreak with Agentic System

Saikat Barua, Mostafizur Rahman, Md Jafor Sadek, Rafiul Islam, Shehenaz Khaled, Ahmedul Kabir

专题命中 越狱攻击 :jailbreak(title);alignment(abstract);safety(abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20099 2025-06-05 cs.CR 82%

Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks

Chen Xiong, Xiangyu Qi, Pin-Yu Chen, Tsung-Yi Ho

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06426 2025-05-30 cs.CR cs.AI cs.CL cs.LG 82%

SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains

Bijoy Ahmed Saiem, MD Sadik Hossain Shanto, Rakib Ahsan, Md Rafi ur Rashid

机构 * Bangladesh University of Engineering and Technology(孟加拉工程科技大学) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12443 2025-05-20 cs.RO 82%

BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation

Wenqi Lyu, Zerui Li, Yanyuan Qiao, Qi Wu

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15207 2025-04-02 cs.SE 82%

Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks

Shide Zhou, Tianlin Li, Kailong Wang, Yihao Huang, Ling Shi, Yang Liu, Haoyu Wang

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01229 2025-03-04 cs.LG cs.AI cs.CL cs.CR math.OC 82%

Boosting Jailbreak Attack with Momentum

Yihao Zhang, Zeming Wei

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05223 2025-02-11 cs.CR cs.AI cs.CL cs.LG 82%

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Buyun Liang, Kwan Ho Ryan Chan, Darshan Thaker, Jinqi Luo, René Vidal

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02705 2025-02-06 cs.CL cs.AI cs.CR cs.LG 82%

Certifying LLM Safety against Adversarial Prompting

Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, Himabindu Lakkaraju

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at COLM 2024: https://openreview.net/forum?id=9Ik05cycLq

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12210 2025-01-22 cs.CR 82%

You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense

Wuyuao Mai, Geng Hong, Pei Chen, Xudong Pan, Baojun Liu, Yuan Zhang, Haixin Duan, Min Yang

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏