WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
WikiSeeker: 重新思考视觉语言模型在基于知识的视觉问答中的作用
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室 (MAIS))
专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(title,abstract);VLM(abstract);分类 cs.CV
AI总结 本文提出WikiSeeker框架,通过引入多模态检索器和重新定义视觉语言模型的角色,提升多模态检索性能和答案质量,实现在EVQA、InfoSeek和M2KR数据集上的最优表现。
Comments Accepted by ACL 2026 Findings