arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3129 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3129 篇

2002.11894 2020-11-24 cs.CV 57%

Unshuffling Data for Improved Generalization

Damien Teney, Ehsan Abbasnejad, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04264 2020-11-10 cs.CL cs.CV 57%

CapWAP: Captioning with a Purpose

Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark, Regina Barzilay

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13128 2020-10-27 cs.AI cs.CL cs.IR 57%

ExplanationLP: Abductive Reasoning for Explainable Science Question Answering

Mokanarangan Thayaparan, Marco Valentino, André Freitas

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11701 2020-10-23 cs.CV cs.CL 57%

Spatial Attention as an Interface for Image Captioning Models

Philipp Sadler

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments A thesis submitted in fulfillment of the requirements for the degree Master of Science in Cognitive Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10604 2020-10-22 stat.ML cs.LG cs.NE 57%

Bayesian Attention Modules

Xinjie Fan, Shujian Zhang, Bo Chen, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.08189 2020-10-19 cs.CV 57%

New Ideas and Trends in Deep Multimodal Content Understanding: A Review

Wei Chen, Weiping Wang, Li Liu, Michael S. Lew

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06572 2020-10-14 cs.CL cs.CV 57%

Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!

Jack Hessel, Lillian Lee

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Journal ref Published in EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14451 2020-10-07 cs.CL cs.CV 57%

Pragmatic Issue-Sensitive Image Captioning

Allen Nie, Reuben Cohn-Gordon, Christopher Potts

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 15 pages, 7 figures. EMNLP 2020 Findings Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00839 2020-10-05 cs.CV 57%

CAPTION: Correction by Analyses, POS-Tagging and Interpretation of Objects using only Nouns

Leonardo Anjoletto Ferreira, Douglas De Rizzo Meneghetti, Paulo Eduardo Santos

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Published at the First Annual International Workshop on Interpretability: Methodologies and algorithms (IMA 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06884 2020-10-05 cs.CV cs.CL 57%

DeVLBert: Learning Deconfounded Visio-Linguistic Representations

Shengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang, Zhou Zhao, Jianke Zhu, Jin Yu, Hongxia Yang, Fei Wu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 10 pages, 4 figures, to appear in ACM MM 2020 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01523 2020-09-04 cs.CV 57%

A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports

Yikuan Li, Hanyin Wang, Yuan Luo

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 10 pages, 3 figures, submitted to BIBM2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10797 2020-08-25 cs.CV 57%

Why do These Match? Explaining the Behavior of Image Similarity Models

Bryan A. Plummer, Mariya I. Vasileva, Vitali Petsiuk, Kate Saenko, David Forsyth

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted at ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.15631 2020-06-30 cs.CV 57%

Improving VQA and its Explanations \\ by Comparing Competing Explanations

Jialin Wu, Liyan Chen, Raymond J. Mooney

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.06637 2020-06-17 cs.CV cs.CL 57%

Exploring Weaknesses of VQA Models through Attribution Driven Insights

Shaunak Halbe

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Second Grand-Challenge and Workshop on Multimodal Language, ACL 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.08322 2020-06-16 cs.CV cs.MM 57%

ORD: Object Relationship Discovery for Visual Dialogue Generation

Ziwei Wang, Zi Huang, Yadan Luo, Huimin Lu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.11872 2020-04-01 cs.CV 57%

Vision and Language: from Visual Perception to Content Creation

Tao Mei, Wei Zhang, Ting Yao

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Journal ref APSIPA Transactions on Signal and Information Processing 9 (2020) e11

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.14080 2020-04-01 cs.CV 57%

X-Linear Attention Networks for Image Captioning

Yingwei Pan, Ting Yao, Yehao Li, Tao Mei

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments CVPR 2020; The source code and model are publicly available at: https://github.com/Panda-Peter/image-captioning

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.12511 2020-03-31 cs.CV 57%

Assessing Image Quality Issues for Real-World Problems

Tai-Yin Chiu, Yinan Zhao, Danna Gurari

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08897 2020-03-20 cs.CV cs.CL cs.MM 57%

Normalized and Geometry-Aware Self-Attention Network for Image Captioning

Longteng Guo, Jing Liu, Xinxin Zhu, Peng Yao, Shichen Lu, Hanqing Lu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09844 2020-03-13 cs.CV 57%

Adversarial Multimodal Network for Movie Question Answering

Zhaoquan Yuan, Siyuan Sun, Lixin Duan, Xiao Wu, Changsheng Xu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments We will revise the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.04732 2020-01-15 cs.CV 57%

Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features

Andres Mafla, Sounak Dey, Ali Furkan Biten, Lluis Gomez, Dimosthenis Karatzas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Winter Conference on Applications of Computer Vision (WACV 2020) Accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.01119 2019-12-10 cs.CV cs.CL 57%

Deep Bayesian Active Learning for Multiple Correct Outputs

Khaled Jedoui, Ranjay Krishna, Michael Bernstein, Li Fei-Fei

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11059 2019-12-05 cs.CV 57%

Unified Vision-Language Pre-Training for Image Captioning and VQA

Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, Jianfeng Gao

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments AAAI 2020 camera-ready version. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07251 2019-11-19 cs.CV cs.CL 57%

DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, Qi Wu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by the Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06352 2019-11-18 cs.CV cs.CL 57%

Question-Conditioned Counterfactual Image Generation for VQA

Jingjing Pan, Yash Goyal, Stefan Lee

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by the VQA Workshop at CVPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03083 2019-11-11 cs.CV cs.CL 57%

Are we asking the right questions in MovieQA?

Bhavan Jasani, Rohit Girdhar, Deva Ramanan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Spotlight presentation at CLVL workshop, ICCV 2019. Project page: https://bhavanj.github.io/MovieQAWithoutMovies/

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.11532 2019-11-11 cs.LG stat.ML 57%

Structure Learning for Neural Module Networks

Vardaan Pahuja, Jie Fu, Sarath Chandar, Christopher J. Pal

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.03285 2019-09-24 cs.CY cs.CV cs.HC 57%

Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image Retrieval

Arijit Ray, Yi Yao, Rakesh Kumar, Ajay Divakaran, Giedrius Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 2019 AAAI Conference on Human Computation and Crowdsourcing

Journal ref 2019 AAAI Conference on Human Computation and Crowdsourcing

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.00313 2019-08-27 cs.CV 57%

VrR-VG: Refocusing Visually-Relevant Relationships

Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, Tao Mei

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by ICCV2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05067 2019-08-15 cs.CL cs.CV 57%

Reactive Multi-Stage Feature Fusion for Multimodal Dialogue Modeling

Yi-Ting Yeh, Tzu-Chuan Lin, Hsiao-Hua Cheng, Yu-Hsuan Deng, Shang-Yu Su, Yun-Nung Chen

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted for a poster session at the DSTC7 workshop at AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏