When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
当提示变为视觉:面向大型图像编辑模型的以视觉为中心的越狱攻击
专题命中 图像编辑 :image editing(title,abstract);分类 cs.CV
AI总结 提出首个视觉到视觉的越狱攻击VJA,通过纯视觉输入传递恶意指令,并构建安全基准IESBench,在商业模型上攻击成功率高达80.9%,同时提出无需训练的内省多模态推理防御方法。
Comments Accepted for spotlight and oral presentation at ICML 2026 (Project: https://csu-jpg.github.io/vja.github.io/)