Defending against Indirect Prompt Injection by Instruction Detection
对抗间接提示注入的指令检测
机构 * Renmin University of China(中国人民大学) ; Peking University Shenzhen Graduate School(北京大学深圳研究生院) ; Wuhan University(武汉大学) ; University of Science and Technology of China(中国科学技术大学) ; Hong Kong University of Science and Technology(香港科技大学) ; Sony AI(索尼人工智能) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI
AI总结 本文提出InstructDetector,通过检测LLMs行为状态来识别IPI攻击,实现高检测准确率和低攻击成功率。
Comments 16 pages, 4 figures