Evaluation of Prompt Injection Defenses in Large Language Models
对大型语言模型中提示注入防御措施的评估
机构 * Swept AI ; University of Michigan(密歇根大学)
专题命中 提示注入 :prompt injection(title);分类 cs.AI
AI总结 研究通过构建自适应攻击者测试了九种防御配置,发现依赖模型自我保护的防御均失效,而输出过滤通过硬编码规则在应用代码中检查响应,实现了零泄露。
Comments 14 pages, 9 figures