arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 528 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 528 篇

2506.09956 2025-06-12 cs.CR cs.AI 79%

LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge

Sahar Abdelnabi, Aideen Fay, Ahmed Salem, Egor Zverev, Kai-Chieh Liao, Chi-Huang Liu, Chun-Chih Kuo, Jannis Weigend, Danyael Manlangit, Alex Apostolov, Haris Umair, João Donato, Masayuki Kawakita, Athar Mahboob, Tran Huu Bach, Tsun-Han Chiang, Myeongjin Cho, Hajin Choi, Byeonghyeon Kim, Hyeonjin Lee, Benjamin Pannell, Conor McCauley, Mark Russinovich, Andrew Paverd, Giovanni Cherubin

机构 * Microsoft(微软) ISTA Trend Micro RainaResearch University of Coimbra(科英布拉大学) Vietnamese German University(越南-德国大学) SK Shieldus HiddenLayer

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Dataset at: https://huggingface.co/datasets/microsoft/llmail-inject-challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05174 2025-06-12 cs.CR cs.AI 79%

MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents

Kaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo, William Yang Wang

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08351 2025-06-10 cs.CL 79%

Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring

Mina Almasi, Ross Deans Kristensen-McLachlan

机构 * Department of Linguistics, Cognitive Science, and Semiotics(语言学、认知科学与符号学系)

专题命中 提示注入 :alignment(title,abstract);分类 cs.CL

Comments Accepted at BEA2025 (Conference workshop at ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05739 2025-06-09 cs.CR cs.AI 79%

To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt

Zhilong Wang, Neha Nagaraja, Lan Zhang, Hayretdin Bahsi, Pawan Patil, Peng Liu

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments To appear in the Industry Track of the 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05446 2025-06-09 cs.CR cs.AI 79%

Sentinel: SOTA model to protect against prompt injections

Dror Ivry, Oran Nahum

机构 * Qualifire

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18333 2025-05-27 cs.CR cs.AI 79%

A Critical Evaluation of Defenses against Prompt Injection Attacks

Yuqi Jia, Zedian Shao, Yupei Liu, Jinyuan Jia, Dawn Song, Neil Zhenqiang Gong

机构 * Duke University(杜克大学) The Pennsylvania State University(宾夕法尼亚州立大学) UC Berkeley(伯克利大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18575 2025-05-20 cs.CR cs.AI 79%

WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, Kamalika Chaudhuri

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Code and data: https://github.com/facebookresearch/wasp

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09798 2025-05-13 cs.CR cs.CL 79%

Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface

Andrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu, Earlence Fernandes

机构 * UC San Diego(加州大学圣地亚哥分校) University of Wisconsin Madison(威斯康星大学麦迪逊分校)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

Journal ref Proceedings of the 2025 IEEE Symposium on Security and Privacy, IEEE Computer Society, 2025, pp. 374-392

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18333 2025-04-28 cs.CR cs.CL 79%

Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections

Narek Maloyan, Dmitry Namiot

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14729 2025-04-07 cs.CR cs.AI 79%

PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs

Jiahao Yu, Yangguang Shao, Hanwen Miao, Junzheng Shi

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00061 2025-03-05 cs.CR cs.LG 79%

Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents

Qiusi Zhan, Richard Fang, Henil Shalin Panchal, Daniel Kang

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

Comments 17 pages, 5 figures, 6 tables (NAACL 2025 Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08966 2025-02-17 cs.CR cs.AI 79%

RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage

Peter Yong Zhong, Siyuan Chen, Ruiqi Wang, McKenna McCall, Ben L. Titzer, Heather Miller, Phillip B. Gibbons

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21492 2024-11-27 cs.CR cs.CL 79%

FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks

Jiongxiao Wang, Fangzhou Wu, Wendi Li, Jinsheng Pan, Edward Suh, Z. Morley Mao, Muhao Chen, Chaowei Xiao

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13352 2024-11-26 cs.CR cs.LG 79%

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, Florian Tramèr

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

Comments Updated version after fixing a bug in the Llama implementation and updating the travel suite

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05047 2024-11-20 cs.CL 79%

A test suite of prompt injection attacks for LLM-based machine translation

Antonio Valerio Miceli-Barone, Zhifan Sun

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20911 2024-11-19 cs.CR cs.AI 79%

Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks

Dario Pasquini, Evgenios M. Kornaropoulos, Giuseppe Ateniese

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments v0.2 (evaluated on more agents)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22284 2024-10-30 cs.CR cs.LG 79%

Embedding-based classifiers can detect prompt injection attacks

Md. Ahsan Ayub, Subhabrata Majumdar

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14479 2024-10-21 cs.CR cs.LG 79%

Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models

Cody Clop, Yannick Teglia

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08776 2024-10-15 cs.CR cs.AI 79%

F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

Yupeng Ren

专题命中 提示注入 :prompt injection(title);safety(abstract);分类 cs.AI

Comments 1. Fixed typo in abstract 2. Provisionally completed the article update to facilitate future version revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13331 2024-09-23 cs.CL cs.CR 79%

Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection

Md Abdur Rahman, Hossain Shahriar, Fan Wu, Alfredo Cuzzocrea

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02691 2024-08-06 cs.CL cs.CR 79%

InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Qiusi Zhan, Zhixiang Liang, Zifan Ying, Daniel Kang

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

Comments 36 pages, 6 figures, 13 tables (ACL 2024 Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09832 2024-03-18 cs.CL 79%

Scaling Behavior of Machine Translation with Large Language Models under Prompt Injection Attacks

Zhifan Sun, Antonio Valerio Miceli-Barone

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL

Comments 15 pages, 18 figures, First Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04957 2024-03-11 cs.AI 79%

Automatic and Universal Prompt Injection Attacks against Large Language Models

Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, Chaowei Xiao

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Pre-print, code is available at https://github.com/SheltonLiu-N/Universal-Prompt-Injection

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07612 2024-01-17 cs.CR cs.AI 79%

Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications

Xuchen Suo

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01011 2023-11-03 cs.LG cs.CR 79%

Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game

Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, Stuart Russell

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06363 2024-09-27 cs.CR 79%

StruQ: Defending Against Prompt Injection with Structured Queries

Sizhe Chen, Julien Piet, Chawin Sitawarin, David Wagner

专题命中 提示注入 :prompt injection(title,abstract)

Comments To appear at USENIX Security Symposium 2025. Key words: prompt injection defense, LLM security, LLM-integrated applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18312 2024-10-25 cs.CR cs.AI cs.CY 79%

Countering Autonomous Cyber Threats

Kade M. Heckel, Adrian Weller

专题命中 提示注入 :safety(abstract);prompt injection(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 76 pages, MPhil Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08417 2026-08-13 cs.CR 版本更新 78%

Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs

注意力是所有你需要的:用于防御大语言模型中的间接提示注入攻击

Yinan Zhong, Qianhao Miao, Yanjiao Chen, Jiangyi Deng, Yushi Cheng, Wenyuan Xu

专题命中 提示注入 :prompt injection(title,abstract)

AI总结 本文提出Rennervate框架,通过注意力机制在token级别检测并清除IPI攻击,有效提升LLM的安全性。

Comments Accepted by Network and Distributed System Security (NDSS) Symposium 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05495 2026-08-07 cs.CR cs.HC 新提交 78%

PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents

PromptShield Home:面向智能家居智能体的环境多模态提示注入防御方案

He Zhang, Feilong Li, Dingning Long, Yilin Cui, Peijun Zhang, Yuewen Zhang, Qianyao Xu, Xinyi Fu

专题命中 提示注入 :prompt injection(title);safety(abstract)

AI总结 该研究针对智能家居智能体的提示注入安全问题,构建了现实场景基准测试PromptShield-Home,对比了三类防御层的性能,发现多智能体调解的上限性能最优,提出应采用学习路由与传感器融合方案保障安全。

Comments This work has been accepted as a poster to UbiComp 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28165 2026-08-03 cs.CR 版本更新 78%

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

利用感知能力:针对多模态大语言模型智能体的隐蔽并发音频提示注入攻击

Mingxiao Liu, Yitong Li, Haoren Zhao, Yaoxiang Bian, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang

专题命中 提示注入 :prompt injection(title,abstract)

AI总结 该研究针对多模态LLM智能体提出隐蔽并发音频提示注入攻击,构建首个相关基准并评估多款智能体,还提出CADV防御机制,经实验验证攻击有效且防御可靠。

Comments 19 pages, 8 figures, The code is publicly available at https://github.com/Limax666/AudioAgentSecurity

详情

展开后加载摘要…

URL PDF HTML 收藏