arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1715 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1715 篇

2507.05248 2025-11-24 cs.CL 79%

Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models

响应攻击:利用上下文优先级来劫持大语言模型

Ziqi Miao, Lijun Li, Yuan Xiong, Zhenhua Liu, Pengyu Zhu, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Xi’an Jiaotong University(西安交通大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

AI总结 响应攻击通过利用对话中中间响应作为上下文优先级,有效劫持大语言模型生成违反政策的内容。

Comments 20 pages, 10 figures. Code and data available at https://github.com/Dtc7w3PQ/Response-Attack

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06194 2025-11-18 cs.CL 79%

SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation

Lai Jiang, Yuekang Li, Xiaohan Zhang, Youtao Ding, Li Pan

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

Comments This paper has been accepted by AAAI 2026 as a poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07480 2025-11-12 cs.CR cs.AI 79%

KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs

Shuyuan Liu, Jiawei Chen, Xiao Yang, Hang Su, Zhaoxia Yin

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03271 2025-11-06 cs.CR cs.CL 79%

Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs

Yize Liu, Yunyun Hou, Aina Sui

专题命中 越狱攻击 :jailbreak(title);red teaming(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00556 2025-11-04 cs.CL 79%

Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack

Peng Ding, Jun Kuang, Wen Sun, Zongyu Wang, Xuezhi Cao, Xunliang Cai, Jiajun Chen, Shujian Huang

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Meituan Inc.(美团公司)

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL

Comments Preprint, 14 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18728 2025-10-22 cs.CR cs.AI 79%

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

Sidhant Narula, Javad Rafiei Asl, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi

机构 * University of Arizona(亚利桑那大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments This paper has been accepted for presentation at the Conference on Applied Machine Learning in Information Security (CAMLIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09023 2025-10-13 cs.LG cs.CR 79%

The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

Milad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff, Jamie Hayes, Michael Ilie, Juliette Pluto, Shuang Song, Harsh Chaudhari, Ilia Shumailov, Abhradeep Thakurta, Kai Yuanqing Xiao, Andreas Terzis, Florian Tramèr

机构 * OpenAI Anthropic Google DeepMind(谷歌DeepMind) HackAPrompt Northeastern University(东北大学) ETH Zürich(苏黎世联邦理工学院) AI Sequrity Company(AI安全公司) MATS Main contributors(MATS主要贡献者)

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06594 2025-10-13 cs.CL 79%

Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?

Sri Durga Sai Sowmya Kadali, Evangelos E. Papalexakis

机构 * Dept. of Computer Science and Engineering University of California, Riverside(计算机科学与工程系加州大学河滨分校)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20848 2025-08-29 cs.CR cs.AI 79%

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring

Junjie Chu, Mingjie Li, Ziqing Yang, Ye Leng, Chenhao Lin, Chao Shen, Michael Backes, Yun Shen, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) Xi’an Jiaotong University(西安交通大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments 17 pages, 5 figures. For the code and data supporting this work, see https://trustairlab.github.io/jades.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05934 2025-08-19 cs.CR cs.AI 79%

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

Ma Teng, Jia Xiaojun, Duan Ranjie, Li Xinfeng, Huang Yihao, Jia Xiaoshuang, Chu Zhixuan, Ren Wenqi

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03054 2025-08-06 cs.AI 79%

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning

Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang

机构 * Key Laboratory of Trustworthy Distributed Computing and Service (MoE)(可信分布式计算与服务重点实验室)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02581 2025-07-25 cs.AI 79%

Neurodivergent Influenceability as a Contingent Solution to the AI Alignment Problem

Alberto Hernández-Espinosa, Felipe S. Abrahão, Olaf Witkowski, Hector Zenil

机构 * Oxford Immune Algorithmics(牛津免疫算法公司) Oxford University Innovation(牛津大学创新) London Institute for Healthcare Engineering(伦敦医疗工程研究所) University of Tokyo(东京大学) The Arrival Institute(抵达研究所) Cancer Research Group(癌症研究组) The Francis Crick Institute(弗朗西斯·克里克研究所) Cross Labs(交叉实验室) King’s Institute for Artificial Intelligence(国王人工智能研究所) King’s College London(伦敦国王学院) The Alan Turing Institute(艾伦·图灵研究所)

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.AI

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09164 2025-07-23 cs.CR cs.AI 79%

ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs

Yuchen Yang, Yiming Li, Hongwei Yao, Bingrun Yang, Yiling He, Tianwei Zhang, Dacheng Tao, Zhan Qin

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou(杭州高新区(滨江)区块链与数据安全研究院,杭州) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08862 2025-07-15 cs.CR cs.CL 79%

RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation

Tianzhe Zhao, Jiaoyan Chen, Yanchi Ru, Haiping Zhu, Nan Hu, Jun Liu, Qika Lin

机构 * School of Computer Science and Technology,Xi'an Jiaotong University(西安交通大学计算机科学与技术学院) Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) National University of Singapore(新加坡国立大学)

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00841 2025-07-02 cs.AI cs.CR 79%

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents

Siyuan Liang, Tianmeng Fang, Zhe Liu, Aishan Liu, Yan Xiao, Jinyuan He, Ee-Chien Chang, Xiaochun Cao

机构 * Nanyang Technological University(南洋理工大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Beihang University(北京航空航天大学) Sun Yat-sen University(中山大学) Beijing Institute of Technology(北京理工大学) National University of Singapore(新加坡国立大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16155 2025-06-27 cs.CL 79%

A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns

Tianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与复杂系统决策智能重点实验室,自动化研究所,中国科学院,北京,中国) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10022 2025-06-13 cs.CR cs.AI 79%

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges

Haoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Beihang University(北京航空航天大学) Beijing Jiaotong University(北京交通大学) Northwestern Polytechnical University(西北工业大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments Accepted as ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14316 2025-05-21 cs.CR cs.AI 79%

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion

Tiehan Cui, Yanxu Mao, Peipei Liu, Congying Liu, Datao You

机构 * Zhongguancun Laboratory(中关村实验室) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00735 2025-05-20 cs.CR cs.AI cs.SE 79%

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

Chun Wai Chiu, Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo

机构 * School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院) University of Chicago(芝加哥大学) University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20493 2025-04-30 cs.CR cs.AI 79%

Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression

Yu Cui, Yujun Cai, Yiwei Wang

机构 * University of California, Merced(加州大学梅尔德分校) University of Queensland(昆士兰大学)

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16489 2025-04-24 cs.CR cs.AI 79%

Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

Senmao Qi, Yifei Zou, Peng Li, Ziyi Lin, Xiuzhen Cheng, Dongxiao Yu

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments 33 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19672 2025-02-28 cs.CV cs.LG 79%

Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack

Chenhe Gu, Jindong Gu, Andong Hua, Yao Qin

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.LG

Comments arXiv admin note: text overlap with arXiv:2403.09766

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13148 2025-02-25 cs.LG cs.CR 79%

Defending Jailbreak Prompts via In-Context Adversarial Game

Yujun Zhou, Yufei Han, Haomin Zhuang, Kehan Guo, Zhenwen Liang, Hongyan Bao, Xiangliang Zhang

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.LG

Comments EMNLP 2024 Main Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08754 2025-02-19 cs.CL cs.CR 79%

StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures

Bangxin Li, Hengrui Xing, Cong Tian, Chao Huang, Jin Qian, Huangqing Xiao, Linfeng Feng

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09113 2025-02-13 cs.LG 79%

Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization

Kai Hu, Weichen Yu, Yining Li, Kai Chen, Tianjun Yao, Xiang Li, Wenhe Liu, Lijun Yu, Zhiqiang Shen, Matt Fredrikson

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.LG

Journal ref 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02438 2025-02-05 cs.CR cs.AI 79%

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

Yaling Shen, Zhixiong Zhuang, Kun Yuan, Maria-Irina Nicolae, Nassir Navab, Nicolas Padoy, Mario Fritz

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.AI

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17522 2025-01-07 cs.CL 79%

DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak

Hao Wang, Hao Li, Junda Zhu, Xinyuan Wang, Chengwei Pan, MinLie Huang, Lei Sha

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07672 2024-12-11 cs.CR cs.CL 79%

FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks

Bocheng Chen, Hanqing Guo, Qiben Yan

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05346 2024-12-10 cs.CR cs.LG 79%

BadGPT-4o: stripping safety finetuning from GPT models

Ekaterina Krupkina, Dmitrii Volkov

专题命中 越狱攻击 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10794 2024-12-04 cs.CL 79%

Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

Yuping Lin, Pengfei He, Han Xu, Yue Xing, Makoto Yamada, Hui Liu, Jiliang Tang

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏