Multimodal Rationales for Explainable Visual Question Answering
机构 * University of Twente(特文特大学) ; University of Bath(巴斯大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Accepted to CVPR workshops 2025
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * University of Twente(特文特大学) ; University of Bath(巴斯大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Accepted to CVPR workshops 2025
专题命中 视觉问答 :visual question answering(title);vision language model(abstract);分类 cs.CV
Comments Accepted in ECML PKDD 2025
机构 * Alibaba Group(阿里巴巴集团)
专题命中 视觉问答 :vision language model(title,abstract);分类 cs.CV
Comments 26 pages, 21 figures
专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.AI
Comments Accepted as ACL 2025 Findings
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted at IIAI AAI 2025, the 3rd International Conference on Computational and Data Sciences in Economics and Finance
机构 * The University of Sydney(悉尼大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
机构 * The Department of Electrical and Electronics Engineering(电气与电子工程系) ; Universidad de los Andes(亚利桑德大学)
专题命中 视觉问答 :grounding(title,abstract);分类 cs.CV
机构 * LIPADE, Université Paris Cité, France(巴黎cité大学LIPADE研究所) ; ONERA, France(法国ONERA研究所)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments EARTHVISION 2025 8 pages, 1 page of supplementary material, 4 figures
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments To be published in 2025 6th International Conference on Computer Vision and Computational Intelligence (CVCI 2025)
专题命中 视觉问答 :grounding(title,abstract);分类 cs.CV
Comments Accepted by IEEE TMM
机构 * University of Edinburgh(爱丁堡大学) ; NVIDIA(英伟达)
专题命中 视觉问答 :visual question answering(title);multimodal large language model(abstract);分类 cs.CV
机构 * Technical University of Munich(慕尼黑技术大学)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted for publication at MIDL 2025
机构 * School of Computer Science, Nanjing University of Information Science & Technology(信息科学技术南京大学计算机科学学院) ; Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET)(江苏省大气环境与装备技术协同创新中心) ; School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学) ; Faculty of Information Technology and Communication Sciences, Tampere University(信息科技与通信科学学院,坦佩雷大学)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by ICME2025
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.AI
专题命中 视觉问答 :MLLM(title,abstract);分类 cs.CV
Comments CVPR 2025
专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV
Comments ICLR 2025 Poster
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2024
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments AAAI-25
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
Comments Preprint
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted to NAACL 2025 Main Conference
专题命中 视觉问答 :vision language model(title,abstract);分类 cs.CV
Comments 19 pages, 10 figures, Accepted to NAACL 2025
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Accepted to ICME2024
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI
Comments 15 pages, 6 figures, 9 tables