AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
AttnTrace: 基于提示注入和知识腐败的上下文归因
机构 * The Pennsylvania State University(宾夕法尼亚州立大学)
专题命中 记忆与上下文管理 :autonomous agent(abstract);分类 cs.CL
AI总结 本文提出AttnTrace,一种基于LLM对提示生成的注意力权重的上下文追溯方法,通过增强效果和技术改进,提高了准确性和效率,并展示了其在检测长上下文中的提示注入应用。
Comments To appear in IEEE S&P 2026. The code is available at https://github.com/Wang-Yanting/AttnTrace. The demo is available at https://huggingface.co/spaces/SecureLLMSys/AttnTrace