The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails
《狼来了:多模态安全护栏中安全输入被对抗性误分类为不安全》
专题命中 图文多模态 :multimodal(title,abstract)
AI总结 该研究针对多模态安全护栏提出不安全诱导攻击,通过不安全语义蒸馏实现84%攻击成功率,揭示了当前多模态安全架构存在安全输入被误判为不安全的可用性漏洞。
Comments Accepted by KDD 2026