arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 528 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 528 篇

2505.06311 2026-01-07 cs.CR cs.AI 79%

Defending against Indirect Prompt Injection by Instruction Detection

对抗间接提示注入的指令检测

Tongyu Wen, Chenglong Wang, Xiyuan Yang, Haoyu Tang, Yueqi Xie, Lingjuan Lyu, Zhicheng Dou, Fangzhao Wu

机构 * Renmin University of China(中国人民大学) Peking University Shenzhen Graduate School(北京大学深圳研究生院) Wuhan University(武汉大学) University of Science and Technology of China(中国科学技术大学) Hong Kong University of Science and Technology(香港科技大学) Sony AI(索尼人工智能) Microsoft Research Asia(微软亚洲研究院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文提出InstructDetector,通过检测LLMs行为状态来识别IPI攻击,实现高检测准确率和低攻击成功率。

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16307 2025-12-19 cs.CR cs.AI 79%

Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks

超越基准:对抗提示注入攻击的创新防御

Safwan Shaheer, G. M. Refatul Islam, Mohammad Rafid Hamid, Tahsin Zaman Jilan

机构 * Dept. of CS, SDS BRAC University(计算机科学系,BRAC大学) Dept. of CSE, SDS BRAC University(计算机工程系,BRAC大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文提出创新防御机制,通过迭代优化防御提示,有效缓解LLM中的目标劫持漏洞,提升小型开源模型的安全性与部署效率。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14285 2025-12-18 cs.CR cs.LG 79%

A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks

针对提示注入攻击的多智能体LLM防御管道

S M Asif Hossain, Ruksat Khan Shayoni, Mohd Ruhul Ameen, Akif Islam, M. F. Mridha, Jungpil Shin

机构 * School of Computing, Wichita State University, Kansas, USA(威斯康星州立大学计算机学院) College of Engineering and Computer Sciences, Marshall University, Huntington, WV, USA(马歇尔大学工程与计算机科学学院) Department of Computer Science and Engineering, University of Rajshahi, Bangladesh(拉贾沙希大学计算机科学与工程系) Department of Computer Science and Engineering, American International University-Bangladesh, Dhaka, Bangladesh(美国国际大学-孟加拉国计算机科学与工程系) School of Computer Science and Engineering, The University of Aizu, Aizuwakamatsu, Japan(立命馆大学计算机科学与工程学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

AI总结 本文提出了一种多智能体防御框架,通过协调的LLM代理实时检测并中和提示注入攻击,显著提升了安全性和系统功能。

Comments Accepted at the 11th IEEE WIECON-ECE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09772 2025-12-15 cs.CL 79%

DeepSeek's WEIRD Behavior: The cultural alignment of Large Language Models and the effects of prompt language and cultural prompting

DeepSeek的WEIRD行为:大型语言模型的文化契合与提示语言和文化提示的影响

James Luther, Donald Brown

机构 * School of Data Science University of Virginia(数据科学学院 芝加哥大学)

专题命中 提示注入 :alignment(title,abstract);分类 cs.CL

AI总结 研究探讨了大型语言模型在不同文化提示下的对齐行为,发现DeepSeek-V3和GPT-5在美文化下表现突出,而GPT-4在英文化下更接近中国,文化提示可调整这种对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00966 2025-12-02 cs.CR cs.LG 79%

Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis

通过指令遵循意图分析缓解间接提示注入

Mintong Kang, Chong Xiang, Sanjay Kariyappa, Chaowei Xiao, Bo Li, Edward Suh

机构 * NVIDIA University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

AI总结 IntentGuard通过分析指令遵循意图,有效缓解间接提示注入攻击,保持模型性能并降低攻击成功率

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18416 2025-11-25 cs.LG 79%

Exploring Potential Prompt Injection Attacks in Federated Military LLMs and Their Mitigation

探索联邦军事大语言模型中的潜在提示注入攻击及其缓解方法

Youngjoon Lee, Taehyun Park, Yunho Lee, Jinu Gong, Joonhyuk Kang

机构 * Institute of Information & Communications Technology Planning & Evaluation (IITP)-ITRC (Information Technology Research Center)(信息与通信技术规划与评估机构(IITP)-ITRC(信息技术研究中心))

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

AI总结 本文探讨了联邦军事大语言模型中潜在的提示注入攻击问题,提出人机协作框架结合技术和政策措施来缓解相关风险。

Comments Accepted to the 3rd International Workshop on Dataspaces and Digital Twins for Critical Entities and Smart Urban Communities - IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15759 2025-11-21 cs.CR cs.AI 79%

Securing AI Agents Against Prompt Injection Attacks

保护AI代理免受提示注入攻击

Badrinath Ramakrishnan, Akshaya Balaji

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文提出了一种多层防御框架,通过评估RAG系统中的提示注入风险,将攻击成功率降低至8.7%,同时保持高任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00447 2025-11-19 cs.CR cs.AI 79%

DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion

Ruofan Liu, Yun Lin, Zhiyong Huang, Jin Song Dong

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24967 2025-11-17 cs.CR cs.AI 79%

SecInfer: Preventing Prompt Injection via Inference-time Scaling

Yupei Liu, Yanting Wang, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang Gong

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11358 2025-11-13 cs.CR cs.AI 79%

DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks

Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, Neil Zhenqiang Gong

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Distinguished Paper Award in IEEE Symposium on Security and Privacy, 2025. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05831 2025-11-12 cs.CR cs.AI 79%

Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization

Ishaan Verma, Arsheya Yadav

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05797 2025-11-11 cs.CR cs.AI 79%

When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins

Yigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann, Christopher Kruegel, Giovanni Vigna

机构 * University of California, Santa Barbara(加州大学圣巴bara分校)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments At IEEE S&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26328 2025-10-31 cs.LG 79%

Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections

David Schmotz, Sahar Abdelnabi, Maksym Andriushchenko

机构 * ELLIS Institute Tübingen(图宾根ELLIS研究所) MPI for Intelligent Systems Tübingen(图宾根智能系统研究所) AI Center(人工智能中心)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19844 2025-10-24 cs.CR cs.AI 79%

CourtGuard: A Local, Multiagent Prompt Injection Classifier

Isaac Wu, Michael Maslowski

机构 * Isaac Wu Research Fellow(Isaac Wu 研究员) Non-Trivial Ventures(非平凡企业)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16128 2025-10-21 cs.CR cs.CY 79%

Prompt injections as a tool for preserving identity in GAI image descriptions

Kate Glazko, Jennifer Mankoff

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CY

Comments Accepted as a poster to Soups 2025

Journal ref The Twenty-First Symposium on Usable Privacy and Security (SOUPS 2025) Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12252 2025-10-20 cs.CR cs.AI 79%

PromptLocate: Localizing Prompt Injection Attacks

Yuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia, Neil Gong

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments To appear in IEEE Symposium on Security and Privacy, 2026. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13543 2025-10-16 cs.CR cs.AI 79%

In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers

Avihay Cohen

机构 * Avihay Cohen

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 37 pages , 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 79%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10248 2025-09-26 cs.LG 79%

Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications

Janis Keuper

机构 * Institute for Machine Learning and Analytics (IMLA)(机器学习与分析研究所)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14271 2025-09-19 cs.CR cs.LG 79%

Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense: A 2022 Study of GPT-3 and Contemporary Models

Gustavo Sandoval, Denys Fenchenko, Junyao Chen

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10540 2025-09-16 cs.CR cs.AI 79%

EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System

Pavan Reddy, Aditya Sanjay Gujral

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 8 pages content, 1 page references, 2 figures, Published at AAAI Fall Symposium Series 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09912 2025-09-15 cs.CY cs.CR 79%

When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

Changjia Zhu, Junjie Xiong, Renkai Ma, Zhicong Lu, Yao Liu, Lingyao Li

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07617 2025-09-10 cs.AI 79%

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

Minghui Li, Hao Zhang, Yechao Zhang, Wei Wan, Shengshan Hu, pei Xiaobing, Jing Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络安全学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13214 2025-08-20 cs.CR cs.AI 79%

Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions

Xuyang Guo, Zekai Huang, Zhao Song, Jiahao Zhang

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15219 2025-07-22 cs.CR cs.AI 79%

PromptArmor: Simple yet Effective Prompt Injection Defenses

Tianneng Shi, Kaijie Zhu, Zhun Wang, Yuqi Jia, Will Cai, Weida Liang, Haonan Wang, Hend Alzahrani, Joshua Lu, Kenji Kawaguchi, Basel Alomair, Xuandong Zhao, William Yang Wang, Neil Gong, Wenbo Guo, Dawn Song

机构 * UC Berkeley(加州大学伯克利分校) UC Santa Barbara(加州大学圣巴巴拉分校) Duke University(杜克大学) National University of Singapore(新加坡国立大学) King Abdulaziz City for Science and Technology(国王阿卜杜勒阿齐兹城市科学技术学院) University of Washington(华盛顿大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14799 2025-07-22 cs.CR cs.AI 79%

Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree

Sam Johnson, Viet Pham, Thai Le

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments EMNLP 2025 System Demonstrations Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13169 2025-07-18 cs.CR cs.AI 79%

Prompt Injection 2.0: Hybrid AI Threats

Jeremy McHugh, Kristina Šekrst, Jon Cefalu

机构 * Preamble, Inc.(Preamble公司)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08837 2025-06-30 cs.LG cs.CR 79%

Design Patterns for Securing LLM Agents against Prompt Injections

Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu, Edoardo Debenedetti, Daniel Dobos, Daniel Fabian, Marc Fischer, David Froelicher, Kathrin Grosse, Daniel Naeff, Ezinwanne Ozoani, Andrew Paverd, Florian Tramèr, Václav Volhejn

机构 * Invariant Labs IBM EPFL(苏黎世联邦理工学院) ETH Zurich(苏黎世联邦理工学院) Swisscom(瑞士通信) Google(谷歌) ETH AI Center(苏黎世联邦理工学院人工智能中心) AppliedAI Institute for Europe(欧洲应用AI研究院) Microsoft(微软) Lakera

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18813 2025-06-25 cs.CR cs.AI 79%

Defeating Prompt Injections by Design

Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, Florian Tramèr

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Updated version with newer models and link to the code

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05849 2025-06-17 cs.CR cs.AI 79%

AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents

Zhun Wang, Vincent Siu, Zhe Ye, Tianneng Shi, Yuzhou Nie, Xuandong Zhao, Chenguang Wang, Wenbo Guo, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校) University of California, Santa Barbara(加州大学圣芭芭拉分校) Washington University, Saint Louis(华盛顿大学圣路易斯分校)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏