Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
揭示视觉-语言模型的脆弱性:通过纹理约束扰动和跨模态优化的多模态对抗协同
Xiang Fang, Wanlong Fang, Changshuo Wang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
Comments12 pages, 3 figures, accepted at ICMHI 2026, 10th International Conference on Medical and Health Informatics, Kyoto, Japan. To appear in ACM Conference Proceedings
机构
*
School of Intelligence Science and Technology(智能科学与技术学院)
;
State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)
;
Nanjing University(南京大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.AI、cs.LG
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
看见 vs. 相信:评估开源多模态大模型在反直觉场景中的语言偏见
Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
机构
*
Zhejiang University(浙江大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV、cs.AI
CommentsThis paper has been accepted by the International Journal of Computer Vision (IJCV), 2026. The first two authors contributed equally to this work. 28 pages
机构
*
McGill University(麦吉尔大学)
;
Mila - Quebec AI Institute(魁北克人工智能研究所)
;
University of Cambridge(剑桥大学)
;
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学)
;
University of Toronto(多伦多大学)
;
Salesforce