The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection
注入悖论:通过RAG上下文注入在安全训练的LLM推荐中实现品牌级压制
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 长文档RAG :RAG(title,title_cn);分类 cs.CL
AI总结 研究发现在基于RAG的LLM推荐中,安全训练会导致注入提示反而压制目标品牌推荐率,揭示了安全机制可能被逆向利用的风险。
Comments 18 pages, 1 figure, 15 tables. Accepted at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN), a non-archival venue