CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model
专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract)
Comments Preprint. Under review
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract)
Comments Preprint. Under review
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 8 pages, No figures
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted by 2024 5th International Conference on Electronic Communication and Artificial Intelligence
Journal ref Proceedings of the 2024 5th International Conference on Electronic Communication and Artificial Intelligence (ICECAI), 2024, pp. 681-685
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Code is available at https://github.com/cnzzx/VSA
专题命中 视觉问答 :visual question answering(title,abstract);multimodal large language model(abstract)
Comments This paper has been accapted by 2024 IEEE International Conference on Robotics and Automation (ICRA)
专题命中 视觉问答 :vision language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments EMNLP 2024 main
专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted to EMNLP2024 Findings
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 35 pages, 12 figures, accepted for publication at the 18th International Conference on Document Analysis and Recognition, ICDAR 2024
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted as a full paper by the tinyML Research Symposium 2024
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments A full version of the paper will be released soon. The codes are available at https://github.com/bowen-upenn/Multi-Agent-VQA
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 8 pages,3 figures and 1 page appendix; The processed graphs and codes will be avalibale
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments To be published in AAAI 24
专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 14 pages, 14 figures, 3 tables
专题命中 视觉问答 :multimodal large language model(title);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Preprint version, Under Review
专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract)
Comments Accepted to 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2023)
专题命中 视觉问答 :vision language model(title,abstract);visual question answering(abstract)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted to IJCAI 2023
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 30 pages. arXiv admin note: text overlap with arXiv:2104.00926, arXiv:2110.02526, arXiv:2108.02059, arXiv:1908.01801 by other authors
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Journal ref GIT Journal of Engineering and Technology, 2020
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Findings of EACL 2023. Aishwarya, Ivana, Emanuele and Aida had equal first author contributions. Elnaz and Anita had equal contributions. Aida and Aishwarya had equal senior contributions
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments CVPR 2023
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract)
Comments This work is accepted by the AAAI 2023
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted by the 2022 International Joint Conference on Neural Networks (IJCNN 2022)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Code: https://github.com/lalithjets/Surgical_VQA.git
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at ACL 2022