arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 528 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 528 篇

2602.03117 2026-05-08 cs.CR 50%

AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?

AgentDyn: 您的代理安全防御能在现实世界动态环境中部署吗?

Hao Li, Ruoyao Wen, Shanghao Shi, Ning Zhang, Yevgeniy Vorobeychik, Chaowei Xiao

专题命中 提示注入 :prompt injection(abstract)

AI总结 本文揭示现有代理安全基准的三大缺陷,提出AgentDyn基准,包含60个挑战性开放任务和560个注入测试用例,评估十种先进防御,发现现有防御不足或过度防御,亟需更适应动态环境的基准。

Comments 26 Pages, 17 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01186 2026-05-05 cs.CR 50%

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting

Trace:通过终端行为指纹揭示AI攻击代理

Murali Ediga, Sudipta Chattopadhyay

专题命中 提示注入 :prompt injection(abstract)

AI总结 本文提出Trace框架,通过终端命令序列识别AI攻击代理模型家族,利用防御性提示注入策略提取系统提示,提升攻击意图分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06572 2026-05-04 cs.CR 50%

Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem

工具链中的寄生体:对MCP生态系统的大型分析

Shuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang, Yuchong Xie, Yu Guo, Libo Chen, Shenghong Li, Zhi Xue

专题命中 提示注入 :prompt injection(abstract)

AI总结 研究发现MCP生态中存在寄生工具链攻击,攻击者通过恶意指令渗透工具链,窃取隐私数据,揭示MCP缺乏安全隔离机制,需加强防御。

Comments Accepted by IEEE Symposium on Security and Privacy, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23374 2026-04-28 cs.CR 50%

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents

代理中的幽灵:重新定义LLM代理的信息流跟踪

Yuandao Cai, Wensheng Tang, Cheng Wen, Shengchao Qin

专题命中 提示注入 :prompt injection(abstract)

AI总结 本文提出NeuroTaint框架,通过语义证据、因果推理和持久上下文跟踪,解决LLM代理中信息流跟踪问题,优于现有基准方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04288 2026-04-16 cs.CR cs.SE 50%

LLM-Enabled Open-Source Systems in the Wild: An Empirical Study of Vulnerabilities in GitHub Security Advisories

基于大语言模型的开源系统在现实中的应用:对GitHub安全通告中漏洞的实证研究

Fariha Tanjim Shifat, Hariswar Baburaj, Ce Zhou, Jaydeb Sarker, Mia Mohammad Imran

专题命中 提示注入 :prompt injection(abstract)

AI总结 本研究分析了295份GitHub安全通告,发现大多数漏洞映射到已知CWE,但模型中介暴露不足,建议结合CWE和OWASP视角更全面地评估LLM集成系统的漏洞。

Comments The 2nd International Workshop on Large Language Model Supply Chain Analysis (LLMSC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02375 2026-04-06 cs.SE cs.PL 50%

KAIJU: An Executive Kernel for Intent-Gated Execution of LLM Agents

KAIJU:一种意图门控的LLM代理执行内核

Cormac Guerin, Frank Guerin

专题命中 提示注入 :prompt injection(abstract)

AI总结 KAIJU通过解耦LLM推理层与代理执行流程,引入意图门控执行和执行内核,提升LLM代理在复杂任务中的执行效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15228 2026-02-18 cs.SE 50%

An Empirical Study on the Effects of System Prompts in Instruction-Tuned Models for Code Generation

对指令微调模型在代码生成中系统提示效果的实证研究

Zaiyu Cheng, Antonio Mastropaolo

专题命中 提示注入 :alignment(abstract)

AI总结 本研究探讨了系统提示对代码生成模型性能的影响,发现提示具体性与模型规模、编程语言等因素相关,且不同语言对提示敏感性存在差异。

Comments 34 pages, 12 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10498 2026-02-12 cs.CR 50%

When Skills Lie: Hidden-Comment Injection in LLM Agents

当技能欺骗:LLM代理中的隐藏评论注入

Qianli Wang, Boyang Ma, Minghui Xu, Yue Zhang

专题命中 提示注入 :prompt injection(abstract)

AI总结 研究发现LLM代理在处理隐藏评论时存在安全风险,通过防御性提示可阻止恶意指令并揭示潜在威胁。

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15528 2026-01-23 cs.DC cs.CR 50%

Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform

为中小企业保障LLM即服务的安全性:一个分布式聊天机器人部署平台的行业案例研究

Jiazhu Xie, Bowen Li, Heyu Fu, Chong Gao, Ziqi Xu, Fengling Han

专题命中 提示注入 :prompt injection(abstract)

AI总结 本文提出一个开源多租户平台,帮助中小企业低成本安全部署定制LLM聊天机器人,通过分布式集群和加密网络实现资源池化与隔离,同时集成防提示注入攻击机制。

Comments Accepted by AISC 2026

Journal ref Australasian Information Security Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04508 2025-11-07 cs.CR 50%

Large Language Models for Cyber Security

Raunak Somani, Aswani Kumar Cherukuri

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17884 2025-10-24 cs.CR 50%

PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents

Toby Murray

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02736 2025-10-21 cs.OS cs.SE 50%

AgentSight: System-Level Observability for AI Agents Using eBPF

Yusheng Zheng, Yanpeng Hu, Tong Yu, Andi Quinn

专题命中 提示注入 :prompt injection(abstract)

Journal ref PACMI'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12175 2025-08-19 cs.CR 50%

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous

Ben Nassi, Stav Cohen, Or Yair

专题命中 提示注入 :prompt injection(abstract)

Comments https://sites.google.com/view/invitation-is-all-you-need/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05445 2025-07-09 cs.CR 50%

A Systematization of Security Vulnerabilities in Computer Use Agents

Daniel Jones, Giorgio Severi, Martin Pouliot, Gary Lopez, Joris de Gruyter, Santiago Zanella-Beguelin, Justin Song, Blake Bullwinkel, Pamela Cortez, Amanda Minnich

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10465 2025-04-15 cs.CV 50%

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Tao Zhang, Xiangtai Li, Zilong Huang, Yanwei Li, Weixian Lei, Xueqing Deng, Shihao Chen, Shunping Ji, Jiashi Feng

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03095 2025-04-01 cs.SE 50%

TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Siqi Gu, Quanjun Zhang, Kecheng Li, Chunrong Fang, Fangyuan Tian, Liuchuan Zhu, Jianyi Zhou, Zhenyu Chen

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02926 2025-02-28 cs.CR 50%

Demystifying RCE Vulnerabilities in LLM-Integrated Apps

Tong Liu, Zizhuang Deng, Guozhu Meng, Yuekang Li, Kai Chen

专题命中 提示注入 :prompt injection(abstract)

Journal ref Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security (CCS '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02817 2025-01-31 cs.CR 50%

Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications

Stav Cohen, Ron Bitton, Ben Nassi

专题命中 提示注入 :prompt injection(abstract)

Comments Website: https://sites.google.com/view/compromptmized

详情

展开后加载摘要…

URL PDF HTML 收藏