Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
阅读还是忽略?视觉语言模型中排版攻击鲁棒性和文本识别的统一基准
机构 * Turing Inc.(Turing公司) ; The University of Tokyo(东京大学) ; Institute of Science Tokyo(东京科学研究院) ; Tohoku University(东北大学)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
AI总结 研究大型视觉语言模型排版攻击鲁棒性和文本识别问题,引入RIO-VQA任务及RIO-Bench基准,发现目标中心防御的权衡,提出数据驱动防御基线,提升模型在两方面的性能,凸显现有鲁棒性范围与现实需求的错位。
Comments Accepted at ECCV 2026. The project page is available at: https://turingmotors.github.io/rio-vqa/