arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1717 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1717 篇

2511.12423 2025-11-18 cs.CR cs.LG 57%

GRAPHTEXTACK: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNs

Jiaji Ma, Puja Trivedi, Danai Koutra

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12149 2025-11-18 cs.CR cs.AI cs.CV 57%

AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models

Jiayu Li, Yunhan Zhao, Xiang Zheng, Zonghuan Xu, Yige Li, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) City University of Hong Kong(香港城市大学) Singapore Management University(新加坡管理大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06151 2025-11-13 cs.CR cs.AI 57%

Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems

Haowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li, Yuekai Huang, Dandan Wang, Qing Wang

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06212 2025-11-11 cs.CR cs.AI 57%

RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework

Seif Ikbarieh, Kshitiz Aryal, Maanak Gupta

机构 * Department of Computer Science(计算机科学系) Tennessee Tech University(田纳西科技大学) School of Interdisciplinary Informatics(跨学科信息学学院) University of Nebraska Omaha(内布拉斯加大学奥马哈分校)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23558 2025-11-07 cs.CR cs.AI 57%

Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models

Yiqi Yang, Hongye Fu

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03434 2025-11-06 cs.HC cs.AI cs.MA cs.NI cs.SI 57%

Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond

Botao 'Amber' Hu, Helena Rong

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

Comments Submitted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02823 2025-11-05 cs.AI 57%

Optimizing AI Agent Attacks With Synthetic Data

Chloe Loughridge, Paul Colognese, Avery Griffin, Tyler Tracy, Jon Kutasov, Joe Benton

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00943 2025-11-03 cs.CR cs.AI 57%

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

Chloe Li, Mary Phuong, Noah Y. Siegel

机构 * University College London(伦敦大学学院)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Accepted to IJCNLP-AACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03728 2025-10-28 cs.AI cs.HC 57%

PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming

Wesley Hanwen Deng, Sunnie S. Y. Kim, Akshita Jha, Ken Holstein, Motahhare Eslami, Lauren Wilcox, Leon A Gatys

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07153 2025-10-22 cs.CR cs.AI 57%

Mind the Web: The Security of Web Use Agents

Avishag Shapira, Parth Atulbhai Gandhi, Edan Habler, Asaf Shabtai

机构 * Ben-Gurion University of the Negev, Israel(内盖夫本·古里安大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16794 2025-10-21 cs.CR cs.LG 57%

Black-box Optimization of LLM Outputs by Asking for Directions

Jie Zhang, Meng Ding, Yang Liu, Jue Hong, Florian Tramèr

机构 * ETH Zurich(苏黎世联邦理工学院) University at Buffalo(布法罗大学) Bytedance, Security Research(字节跳动安全研究)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14845 2025-10-17 cs.LG cs.CV 57%

Backdoor Unlearning by Linear Task Decomposition

Amel Abdelraheem, Alessandro Favero, Gerome Bovet, Pascal Frossard

机构 * EPFL(苏黎世联邦理工学院) Cyber-Defence Campus, armasuisse(网络安全校园,armasuisse)

专题命中 越狱攻击 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12575 2025-10-14 cs.CR cs.AI 57%

DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent

Pengyu Zhu, Zhenhong Zhou, Yuanhe Zhang, Shilinlu Yan, Kun Wang, Sen Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01444 2025-10-10 cs.CR cs.AI 57%

PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization

Aofan Liu, Lulu Tang, Ting Pan, Yuguo Yin, Bin Wang, Ao Yang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

Comments Accepted to IEEE International Conference on Multimedia and Expo (ICME) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07192 2025-10-09 cs.LG 57%

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies, Burak Hasircioglu, Ezzeldin Shereen, Carlos Mougan, Vasilios Mavroudis, Erik Jones, Chris Hicks, Nicholas Carlini, Yarin Gal, Robert Kirk

机构 * UK AI Security Institute(英国人工智能安全研究所) Anthropic Alan Turing Institute(艾伦·图灵研究所) OATML, University of Oxford(OATML,牛津大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 越狱攻击 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23882 2025-10-07 cs.AI cs.CR 57%

Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B

Shuyi Lin, Tian Lu, Zikai Wang, Bo Wen, Yibo Zhao, Cheng Tan

机构 * Northeastern University(东北大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03204 2025-10-06 cs.CL 57%

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar, Xing Han Lù, Léo Boisvert, Massimo Caccia, Jérémy Espinas, Alexandre Aussem, Véronique Eglin, Alexandre Lacoste

机构 * LIRIS - CNRS, INSA Lyon, Universite Claude Bernard Lyon 1(LIRIS - CNRS,INSA里昂,克劳德·贝尔纳大学里昂) Esker ServiceNow Research(ServiceNow研究) Mila - Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15420 2025-10-01 cs.CR cs.AI 57%

Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries

Yuhao Wang, Wenjie Qu, Shengfang Zhai, Yanze Jiang, Zichen Liu, Yue Liu, Yinpeng Dong, Jiaheng Zhang

机构 * National University of Singapore(新加坡国立大学) Peking University(北京大学) Tsinghua University(清华大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13899 2025-09-30 cs.LG 57%

Causes and Consequences of Representational Similarity in Machine Learning Models

Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo, Emily Wenger

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17371 2025-09-24 cs.CR cs.LG 57%

SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models

Haotian Xu, Qingsong Peng, Jie Shi, Huadi Zheng, Yu Li, Cheng Zhuo

机构 * Zhejiang University(浙江大学) Huawei(华为)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15739 2025-09-22 cs.CL 57%

Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics

Reza Sanayei, Srdjan Vesic, Eduardo Blanco, Mihai Surdeanu

机构 * Department of Computer Science, University of Arizona(亚利桑那大学计算机科学系) CRIL CNRS & University of Artois(CNRS CRIL与阿维尼昂大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02844 2025-09-17 cs.CV cs.CL cs.CR 57%

Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection

Ziqi Miao, Yi Ding, Lijun Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main). 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15244 2025-09-17 cs.CV cs.AI 57%

Adversarial Prompt Distillation for Vision-Language Models

Lin Luo, Xin Wang, Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07941 2025-09-10 cs.CR cs.AI 57%

ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation

Kai Ye, Liangcai Su, Chenxiong Qian

机构 * The University of Hong Kong(香港大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

Comments This paper has been accepted by the ACM Conference on Computer and Communications Security (CCS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19288 2025-08-28 cs.CR cs.AI 57%

Tricking LLM-Based NPCs into Spilling Secrets

Kyohei Shiomi, Zhuotao Lian, Toru Nakanishi, Teruaki Kitasuka

机构 * Hiroshima University(广岛大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17244 2025-08-26 cs.AI 57%

L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems

Aoun E Muhammad, Kin-Choong Yow, Nebojsa Bacanin-Dzakula, Muhammad Attique Khan

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15283 2025-08-22 cs.IR cs.CL 57%

Adversarial Attacks against Neural Ranking Models via In-Context Learning

Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, Charles L. A. Clarke

机构 * University of Waterloo(多伦多大学) University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20841 2025-08-19 cs.CL 57%

Concealment of Intent: A Game-Theoretic Analysis

Xinbo Wu, Abhishek Umrawal, Lav R. Varshney

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02141 2025-08-13 cs.CV cs.CL 57%

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, Linlin Shen

机构 * Shenzhen University(深圳大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) City University of Hong Kong(香港城市大学) Stanford University(斯坦福大学) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments ICCV 2025, 38 pages, 22 figures, 35 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04632 2025-08-08 cs.CL 57%

IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards

Xu Guo, Tianyi Liang, Tong Jian, Xiaogui Yang, Ling-I Wu, Chenhui Li, Zhihui Lu, Qipeng Guo, Kai Chen

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏