When the Defense Writes the Refusal: Auditing Keyword-Scored Evaluation of Inference-Time Defenses for Multimodal Large Language Models
多模态大语言模型推理时防御方法的比较分析
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)
AI总结 本文比较评估了三种推理时防御方法及其组合在InternVL和Qwen-VL系列共8个模型上的效果,发现无单一防御在所有设置中占优,组合防御导致良性查询过度拒绝率达97-100%,而简单安全提示在保持实用性的同时带来适度安全提升。
Comments 15 pages, 3 figures. Conditionally accepted at DAMDID/RCDL 2026; revised after peer review