Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
Robust-LLaVA:大规模鲁棒图像编码器对多模态大语言模型的有效性
机构 * Mohamed bin Zayed University of AI(Mohamed bin Zayed人工智能大学) ; Khalifa University(卡利法大学) ; Michigan State University(密歇根州立大学) ; Australian National University(澳大利亚国立大学)
专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 本文提出利用大规模对抗预训练的图像分类模型替代CLIP编码器,以增强多模态大语言模型对视觉对抗扰动的鲁棒性,在无需额外对抗训练的情况下,在视觉问答、图像描述和越狱攻击任务中取得显著鲁棒性提升。
Comments Accepted at Trustworthy FMs Workshop Trust Before Use: Building Foundation Models that You Can Trust (ICCVW) 2025