PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted by IJCAI 2024
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted by IJCAI 2024
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted to IEEE Intelligent Systems
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted By AAAI 2024
专题命中 视觉问答 :grounding(title);分类 cs.AI
Comments 10 pages, 3 figures, 2 tables, accepted for the Scholarly-QALD challenge at the International Semantic Web Conference (ISWC) 2023
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments EMNLP 2023
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Published in Robotics (Q1, SCI indexed Journal): https://www.mdpi.com/2218-6581/12/4/114
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments 22 pages
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Yuwei Zhang and Chih-Hui Ho contributed equally to this work
专题命中 视觉问答 :grounding(title);分类 cs.CV
Journal ref IJCAI 2022 Workshop on Spatio-Temporal Reasoning and Learning
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Journal ref CVPR'2022 Workshop on Open-Domain Retrieval Under a Multi-Modal Setting
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted to CVPR 2022. Camera-Ready version. Project page: https://simvqa.github.io/
专题命中 视觉问答 :vision-language model(title);分类 cs.CV
Comments Accepted to ACL 2022
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments NAACL 2021 MAI-Workshop. Code available at https://github.com/codexxxl/GraphVQA
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments arXiv admin note: substantial text overlap with arXiv:1910.10706
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Rank 2 in VQA Challenge 2019
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Published in CVPR 2018
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Submitted to Reproducibility in ML Workshop, ICML'18
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Accepted to CVPR 2018
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
专题命中 视觉问答 :visual question answering(title);分类 cs.CV
Comments Submitted to ECCV 2016
通用科学人工智能的实现路径:科学图像的多模态理解
机构 * TIB Leibniz Information Centre for Science and Technology(莱布尼茨科学与技术信息中心(TIB)) ; Eindhoven University of Technology(埃因霍温理工大学) ; Freie Universität Berlin(柏林自由大学) ; International Center for Chemical and Biological Sciences, University of Karachi(卡拉奇大学国际化学与生物科学中心) ; National University of Sciences and Technology(国家科技大学) ; University of Warwick(华威大学) ; PSG College of Technology(PSG理工学院) ; Aalto University(阿尔托大学)
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 该研究提出以科学图像多模态理解构建通用科学AI的路径,依托ALD/E-ImageMiner基准与2026年ICDAR竞赛,明确相关任务能力检验维度及长期研究方向,推动可机器执行的科学视觉知识与可验证多模态科学AI发展。
Comments 15 pages, 1 figure
CARE:用于可靠医学视觉问答的置信感知推理
机构 * Ant Group(蚂蚁集团) ; University of Michigan(密歇根大学) ; City University of Hong Kong(香港城市大学)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对医学多模态大语言模型的置信度校准问题,提出CARE框架,通过双阶段流程优化准确率与校准度,在三个医学VQA基准上取得最优性能,为临床决策支持提供可信基础。
Comments Accepted by MICCAI 2026
PinpointQA: 一个用于室内视频中微小物体中心空间理解的数据集和基准
机构 * Jilin University(吉林大学) ; National Taiwan University(国立台湾大学)
专题命中 视觉问答 :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 本文提出PinpointQA数据集,用于评估多模态大语言模型在室内视频中对微小物体空间定位的精准理解能力,通过四个递进任务验证模型性能,并展示其作为诊断基准和训练数据集的效果。
Comments Accepted at the ECCV 2026 Workshop on Embodied Multimodal Reasoning in Physical Environments (EMR)
Geo-Embed:面向城市理解的统一多模态嵌入
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.LG
AI总结 本文针对现有多模态嵌入模型难以支持异构地理空间任务的问题,提出GeoMEB基准与Geo-Embed模型,在GeoMEB上实现15.3%的相对性能提升,为地理空间嵌入器发展提供方向。
FAU参加2026年ImageCLEF多模态推理任务:鲁棒候选评分与简洁多语种视觉问答
机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学)
专题命中 视觉问答 :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.LG
AI总结 FAU提交的多模态推理系统,未做任务特定模型训练,在2026年ImageCLEF竞赛中,获Visual MCQ第三名、Visual OpenQA第一名,凸显推理工程的实用价值。
Comments 16 pages, 3 figures, 7 tables. CLEF 2026 Working Notes, ImageCLEF 2026 Multimodal Reasoning Task
LongChart VQA:面向具备复杂多图表推理能力的多模态大语言模型的综合基准
机构 * The University of Hong Kong(香港大学)
专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 LongChart VQA是面向具备复杂多图表推理能力的MLLMs的综合基准,该基准含平均6.5张图像与31.2个问题,评估10个SOTA MLLMs发现其准确率随计算复杂性提升下降,为多图表推理研究指明方向。
ORCA:面向无需训练的3D CT视觉令牌压缩的器官质心聚合方法
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
AI总结 针对3D CT视觉令牌压缩的痛点,提出无需训练的即插即用ORCA方法,在多数据集、多任务上实现显著压缩比与速度提升,性能优于现有方法。
FORGE:用于长视频理解的相关几何中的帧正交性
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Center for Signal and Information Processing (CSIP)(信号与信息处理中心 (CSIP)) ; School of Electrical and Computer Engineering(电气与计算机工程学院)
专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.LG
AI总结 研究长视频理解中推理时最大化查询相关信息问题,提出FORGE模型无关方法,通过在预训练嵌入空间引入查询条件几何统一相关性和多样性,实验证明该方法在关键帧选择和问答上有显著提升。
Comments Under Review
深度专家注入用于带有领域特定知识的视网膜视觉语言模型的锚定
机构 * Beijing Institute of Technology, Beijing, China(北京理工大学) ; National University of Singapore, Singapore(新加坡国立大学) ; Tsinghua University, Beijing, China(清华大学) ; The Hong Kong Polytechnic University, Hong Kong(香港理工大学)
专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
AI总结 本文提出EyExIn框架,通过深度专家注入机制提升视网膜VLMs的领域知识嵌入能力,解决感知与推理间隙问题,实现高精度眼科视觉问答。
穿越幻象:一种双路径代理框架用于鲁棒的误导图表问答
机构 * The Hong Kong University of Science and Technology(香港科技大学)
专题命中 视觉问答 :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 本文提出ChartCynics双路径框架,通过分离感知与验证,利用诊断视觉路径和OCR驱动数据路径解决图表误导问题,提升模型鲁棒性。
Comments 10pages, 4 figures