arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-06-11 至 2026-06-11 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 3 篇

2602.05746 2026-06-11 cs.LG cs.AI 版本更新 84%

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

学习注入:通过强化学习实现自动化提示注入

Xin Chen, Jie Zhang, Florian Tramèr

机构 * ETH Zürich(苏黎世联邦理工学院)

专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI、cs.LG

AI总结 提出AutoInject,一种基于强化学习的黑盒框架,自动学习对抗性后缀进行提示注入,在AgentDojo上优于模板攻击和多种自适应攻击,并突破专门防御模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13631 2026-06-11 stat.CO 版本更新 82%

ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections

ProjGuard:通过低维投影实现计算机使用代理的安全监控

Kebin Contreras, Carlos Hinojosa, Jorge Bacca, Bernard Ghanem

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract)

AI总结 ProjGuard通过行为轨迹监控实现计算机使用代理的安全防护,利用轻量级风险信号提前预警潜在危险,结合辅助视觉语言模型进行针对性修正,提升任务完成率并降低安全风险。

Comments The manuscript was submitted under an inappropriate category. In addition, substantial updates and improvements are currently being made to the document. To avoid confusion and ensure that readers access the most accurate version of the work, we request withdrawal of the current manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10749 2026-06-11 cs.CR 版本更新 78%

AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations

AttriGuard: 通过工具调用的因果归因击败LLM代理中的间接提示注入

Yu He, Haozhe Zhu, Yiming Li, Shuo Shao, Hongwei Yao, Zhihao Liu, Zhan Qin

专题命中 提示注入 :prompt injection(title,abstract)

AI总结 针对LLM代理易受间接提示注入攻击的问题,提出基于并行反事实测试的运行时防御AttriGuard,通过因果归因区分用户意图驱动与不可信观察驱动的工具调用,在多个基准上实现0%攻击成功率。

Comments Accepted by USENIX Security 2026

详情

展开后加载摘要…

URL PDF HTML 收藏