Multi-Sourced Compositional Generalization in Visual Question Answering
Chuanhao Li, Wenbo Ye, Zhen Li, Yuwei Wu, Yunde Jia
机构
*
Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国)
;
Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国)
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering
Zhilin Zhang, Jie Wang, Zhanghao Qin, Ruiqi Zhu, Xiaoliang Gong
机构
*
Tandon School of Engineering, New York University(纽约大学工程学院)
;
School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院)
;
College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
Zhou Yu, Xuecheng Ouyang, Zhenwei Shao, Meng Wang, Jun Yu
机构
*
Key Laboratory of Complex Systems Modeling and Simulation, the School of Computer Science, Hangzhou Dianzi University(复杂系统建模与仿真重点实验室、计算机科学学院、杭州电子大学)
;
HDU-ITMO Joint Institute, Hangzhou Dianzi University(杭州电子大学-ITMO联合学院)
;
School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院、合肥工业大学)
;
School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen(智能科学与工程学院、哈尔滨工业大学深圳校区)
CommentsAn extended journal version of our CVPR 2023 paper, which has been accepted at IEEE T-PAMI 2025. The original conference version can be referred to as the v1 version
CommentsThis version was published at EMNLP 2024 Main Conference as a Long Paper (Oral). See the extended version (arXiv:2411.08870) for additional results on QA tasks based on clinical notes and evaluations in the supervised fine-tuning regime