Knowledge-Based Counterfactual Queries for Visual Question Answering
专题命中 视觉问答 :visual question answering(title,abstract)
Journal ref AAAI MAKE 2023
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(title,abstract)
Journal ref AAAI MAKE 2023
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Published in Cognitive Science 2024
专题命中 视觉问答 :visual question answering(title,abstract)
Comments VLSP2022 EVJVQA challenge
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Accepted at LREC-COLING 2024
专题命中 视觉问答 :vision-language model(title,abstract)
Comments Accepted at Humanoids2023
专题命中 视觉问答 :visual question answering(title,abstract)
专题命中 视觉问答 :visual question answering(title);分类 cs.CV、cs.AI、cs.LG
Comments EMNLP 2023
Journal ref EMNLP 2023
专题命中 视觉问答 :visual question answering(title,abstract)
Comments submitted to Elsevier
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Findings of EACL 2023
专题命中 视觉问答 :visual question answering(title,abstract)
Comments ACL 2023
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Accepted at ACL 2023 as a long paper (Findings)
专题命中 视觉问答 :visual question answering(title,abstract)
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Findings of EACL 2023
专题命中 视觉问答 :visual question answering(title,abstract)
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Accepted to appear at the main conference of EMNLP 2022
专题命中 视觉问答 :grounding(title,abstract)
Comments Accepted as Main Conference Long paper at COLING 2022
专题命中 视觉问答 :grounding(title,abstract)
Comments MM22
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Findings of ACL 2022
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Accepted in EMNLP-Findings (2021)
专题命中 视觉问答 :visual question answering(title,abstract)
Comments Accepted to SIGIR'21 as a short paper
专题命中 视觉问答 :visual question answering(title,abstract)
Comments ICMR '21: ACM International Conference on Multimedia Retrieval, Taipei, Taiwan, August 21-24, 2021
专题命中 视觉问答 :visual question answering(title,abstract)
Comments ACL 2019 (7 pages)
专题命中 视觉问答 :visual question answering(title,abstract)
专题命中 视觉问答 :visual question answering(title);分类 cs.CV、cs.AI、cs.LG
Comments Working notes from the ImageCLEF 2019 VQA-Med competition
重排序器所见:面向长文档多模态问答的多方面页面标注
机构 * Emory University(埃默里大学) ; Hippocratic AI(希波克拉底人工智能公司)
专题命中 视觉问答 :VLM(abstract,abstract_cn);visual question answering(abstract);分类 cs.AI
AI总结 本文针对长文档多模态问答的重排序瓶颈,提出含Trident-R与Trident-S组件的Trident模型,通过多方面页面标注提升检索与生成性能,在多数据集上取得显著效果。
VideoVIBE:用于一次性交互式网站生成的基于视频的诊断基准
专题命中 视觉问答 :MLLM(summary_cn,abstract_cn);分类 cs.CV
AI总结 针对现有一次性交互式网站生成质量评估的不足,提出VideoVIBE基准与V2Lens多智能体系统,实验显示V2Lens可提升Video MLLM的评估性能。
ChronoVision:基于潜在状态重构的时序推理
专题命中 视觉问答 :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 针对多模态大语言模型时序推理不足的问题,本文提出ChronoVision框架,引入Vbvr-VQA数据集,实验显示其在相关基准上取得最优性能。
DocTrace:面向可追溯的长文档视觉问答的分层证据图推理方法
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.AI
AI总结 本文提出DocTrace分层框架,将长文档视觉问答建模为显式证据图推理问题,经两阶段训练优化后,在三个基准上优于现有模型,且推理可追溯。
OncoTriad-QA:用于泛癌推理的患者级放射学-病理学-基因组学基准
专题命中 视觉问答 :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.AI
AI总结 本文提出OncoTriad-QA多模态癌症问答基准及OncoVLM模型,实验显示OncoVLM经微调后在泛癌问答任务上优于现有模型,该基准可用于相关模型的训练与评估。
面向医学视觉基础模型的位置感知细粒度表示学习
机构 * The University of British Columbia(不列颠哥伦比亚大学) ; Vector Institute(矢量研究所)
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
AI总结 本研究提出基于位置感知细粒度表示学习的医学视觉基础模型LoFi,构建大规模医学定位数据集MedG,在多项医学视觉任务中性能优于现有模型。