AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
AttriGuard: 通过工具调用的因果归因击败LLM代理中的间接提示注入
专题命中 其他LLM :LLM(title,title_cn)
AI总结 针对LLM代理易受间接提示注入攻击的问题,提出基于并行反事实测试的运行时防御AttriGuard,通过因果归因区分用户意图驱动与不可信观察驱动的工具调用,在多个基准上实现0%攻击成功率。
Comments Accepted by USENIX Security 2026