arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3129 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3129 篇

2111.14725 2021-11-30 cs.CV 57%

Searching the Search Space of Vision Transformer

Minghao Chen, Kan Wu, Bolin Ni, Houwen Peng, Bei Liu, Jianlong Fu, Hongyang Chao, Haibin Ling

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted to NIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00753 2021-11-29 cs.CV 57%

Structured Multimodal Attentions for TextVQA

Chenyu Gao, Qi Zhu, Peng Wang, Hui Li, Yuliang Liu, Anton van den Hengel, Qi Wu

专题命中 视觉问答 :visual reasoning(abstract);分类 cs.CV

Comments winner of TextVQA Challenge 2020, Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10023 2021-11-22 cs.CV 57%

UFO: A UniFied TransfOrmer for Vision-Language Representation Learning

Jianfeng Wang, Xiaowei Hu, Zhe Gan, Zhengyuan Yang, Xiyang Dai, Zicheng Liu, Yumao Lu, Lijuan Wang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.03642 2021-11-08 cs.CL cs.LG 57%

Grounded Graph Decoding Improves Compositional Generalization in Question Answering

Yu Gai, Paras Jain, Wendi Zhang, Joseph E. Gonzalez, Dawn Song, Ion Stoica

专题命中 视觉问答 :grounding(abstract);分类 cs.LG

Comments To be published in Findings of EMNLP 2021. Code available at https://github.com/gaiyu0/cfq

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07391 2021-09-29 cs.CV 57%

Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval

Kazuya Ueki

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by 20th IEEE International Conference on Machine Learning and Applications (ICMLA2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04422 2021-09-10 cs.CV cs.CL 57%

TxT: Crossmodal End-to-End Learning with Transformers

Jan-Martin O. Steitz, Jonas Pfeiffer, Iryna Gurevych, Stefan Roth

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments To appear at the 43rd DAGM German Conference on Pattern Recognition (GCPR) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12756 2021-08-24 cs.CV cs.CL 57%

InfographicVQA

Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, C. V Jawahar

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.03367 2021-07-22 cs.CV 57%

Disentangling 3D Prototypical Networks For Few-Shot Concept Learning

Mihir Prabhudesai, Shamit Lal, Darshan Patil, Hsiao-Yu Tung, Adam W Harley, Katerina Fragkiadaki

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.00926 2021-07-21 cs.CV cs.HC 57%

VisQA: X-raying Vision and Language Reasoning in Transformers

Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05556 2021-07-09 cs.CL cs.CV 57%

Sparse and Structured Visual Attention

Pedro Henrique Martins, Vlad Niculae, Zita Marinho, André Martins

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09141 2021-06-18 cs.CL cs.CV 57%

Probing Image-Language Transformers for Verb Understanding

Lisa Anne Hendricks, Aida Nematzadeh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05251 2021-06-10 cs.LG cs.CL stat.ML 57%

Bayesian Attention Belief Networks

Shujian Zhang, Xinjie Fan, Bo Chen, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

Comments ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07141 2021-05-18 cs.CV 57%

Show Why the Answer is Correct! Towards Explainable AI using Compositional Temporal Attention

Nihar Bendre, Kevin Desai, Peyman Najafirad

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 7 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.01993 2021-05-06 cs.CV 57%

AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss

Yangyang Guo, Liqiang Nie, Zhiyong Cheng, Feng Ji, Ji Zhang, Alberto Del Bimbo

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.06087 2021-04-20 cs.CV 57%

Contrast and Classify: Training Robust VQA Models

Yash Kant, Abhinav Moudgil, Dhruv Batra, Devi Parikh, Harsh Agrawal

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08108 2021-04-19 cs.CV cs.CL 57%

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

Shir Gur, Natalia Neverova, Chris Stauffer, Ser-Nam Lim, Douwe Kiela, Austin Reiter

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03762 2021-04-09 cs.CV cs.CL 57%

Video Question Answering with Phrases via Semantic Roles

Arka Sadhu, Kan Chen, Ram Nevatia

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments NAACL21 Camera Ready including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03656 2021-04-09 cs.CV 57%

How Transferable are Reasoning Patterns in VQA?

Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche, Romain Vuillemot, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05726 2021-04-09 cs.CV cs.CL 57%

Estimating semantic structure for the VQA answer space

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments [WARNING] We want to notice the reader that additional experiments (not in the paper) have shown that using a `random' semantic space performs as much as the proposed semantic loss. This additional result question the effectiveness of our method

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05121 2021-04-08 cs.CV 57%

Roses Are Red, Violets Are Blue... but Should Vqa Expect Them To?

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.00332 2021-04-02 cs.CV 57%

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, Jingjing Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15974 2021-03-31 cs.CV 57%

Domain-robust VQA with diverse datasets and methods but no target labels

Mingda Zhang, Tristan Maidment, Ahmad Diab, Adriana Kovashka, Rebecca Hwa

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments To appear in CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.08981 2021-03-31 cs.CV cs.CL 57%

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2021). Our dataset is available at https://github.com/google-research-datasets/conceptual-12m

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04181 2021-03-09 cs.LG 57%

Contextual Dropout: An Efficient Sample-Dependent Dropout Module

Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, Mingyuan Zhou

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

Journal ref ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06573 2021-01-19 cs.AI 57%

Understanding in Artificial Intelligence

Stefan Maetschke, David Martinez Iraola, Pieter Barnard, Elaheh ShafieiBavani, Peter Zhong, Ying Xu, Antonio Jimeno Yepes

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments 28 pages, 282 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05153 2020-12-10 cs.CV 57%

Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps

Qi Zhu, Chenyu Gao, Peng Wang, Qi Wu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.02951 2020-12-08 cs.CV 57%

FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding

Maryam Rahnemoonfar, Tashnim Chowdhury, Argho Sarkar, Debvrat Varshney, Masoud Yari, Robin Murphy

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04963 2020-12-02 cs.CV 57%

Rephrasing visual questions by specifying the entropy of the answer distribution

Kento Terao, Toru Tamaki, Bisser Raytchev, Kazufumi Kaneda, Shun'ichi Satoh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14759 2020-11-30 cs.AI 57%

Graph-based Heuristic Search for Module Selection Procedure in Neural Module Network

Yuxuan Wu, Hideki Nakayama

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments in Neural Module Network[C]//Proceedings of the Asian Conference on Computer Vision. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11721 2020-11-25 cs.CV 57%

Siamese Tracking with Lingual Object Constraints

Maximilian Filtenborg, Efstratios Gavves, Deepak Gupta

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏