Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack
先定位再中和:梯度引导的令牌抑制对抗视觉提示注入攻击
机构 * School of Advanced Interdisciplinary Sciences, UCAS(UCAS交叉学科研究院) ; School of Electronic, Electrical and Communication Engineering, UCAS(UCAS电子电气与通信工程学院) ; State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室) ; Alibaba Group(阿里巴巴集团) ; School of Computer Science and Technology, UCAS(UCAS计算机科学与技术学院) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; Key Laboratory of Big Data Mining and Knowledge Management, UCAS(UCAS大数据挖掘与知识管理重点实验室)
专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.LG
AI总结 针对多模态大语言模型的视觉提示注入攻击,提出梯度令牌掩码(GTM)方法,通过梯度分析定位关键图像令牌并掩码中和,将攻击成功率降至接近零且计算开销极小。