From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
从提示注入到持久控制:防御智能体框架中的木马后门
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院 Gallagher 学院)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出ClawTrojan基准测试揭示本地智能体框架中的多步木马攻击,并设计DASGuard防御方法,通过扫描控制文本、追溯来源并清除不可信控制内容,实现动态防御。
Comments Code and data are available at https://github.com/RUC-NLPIR/ClawTrojan