Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message
木马提示:通过伪造助手消息来破解对话式多模态模型
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI
AI总结 研究针对对话式多模态模型的安全漏洞,提出木马提示越狱技术,通过伪造对话历史绕过安全机制,在谷歌Gemini-2.0上实验显示其攻击成功率高于现有方法,揭示对话AI安全缺陷,呼吁转变验证范式。
Comments We stopped working on this idea because we realized that there was a paper submitted before it that took almost identical approach