arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-07-31 至 2026-07-31 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 2 篇

2606.04425 2026-07-31 cs.CR cs.AI 版本更新 79%

What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection

如果提示注入从未消失?探索智能体系统中的跨会话存储提示注入

Yuanbo Xie, Wenlei Zhu, Tianyun Liu, Yingjie Zhang, Suchen Liu, Yulin Li, Liya Su, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) AI Sec Lab, Beijing Chaitin Technology Co.,Ltd(北京柴坦科技有限公司AI安全实验室)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本研究引入跨会话存储提示注入,通过持久化状态使提示注入从单会话模型级威胁转变为长期系统级漏洞,并构建了分类法、基准测试和沙箱工具以评估风险。

Comments position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18248 2026-07-31 cs.CR cs.CL 版本更新 79%

Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection

超越模式匹配:七种跨领域技术用于提示注入检测

Thamilvendhan Munirathinam

机构 * Independent Researcher(独立研究者) prompt-shield project(prompt-shield项目)

专题命中 提示注入 :prompt injection(title);alignment(abstract);分类 cs.CL

AI总结 本文提出七种跨领域技术用于提示注入检测,通过引入来自不同领域的机制,如法医学语言学、材料科学疲劳分析、网络安全欺骗技术、生物信息学本地序列比对、经济学机制设计、流行病学谱信号分析和编译器理论中的污点跟踪,以改进现有的提示注入检测方法。

Comments v4 (31 pp, up from 27): adds Sec. 2.4 concurrent-work map (13 papers), Sec. 4.3 marked implemented (d034 ships in v0.7.3), Sec. 5.7 composed-stack adaptive-attack partial run, Sec. 5.8 50-doc held-out benchmark (d027 1.000->0.000 F1 in isolation, composed engine 0.815 F1), Sec. 7 architectural patterns. Repro tag: v0.7.3. Zenodo DOI 10.5281/zenodo.19644135

详情

展开后加载摘要…

URL PDF HTML 收藏