arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26071 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3129 篇

2104.06365 2021-04-14 cs.LG cs.CV 62%

Neuro-Symbolic VQA: A review from the perspective of AGI desiderata

Ian Berlot-Attwell

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03046 2021-04-08 cs.CV cs.LG 62%

Multimodal Continuous Visual Attention Mechanisms

António Farinhas, André F. T. Martins, Pedro M. Q. Aguiar

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02096 2021-04-07 cs.CV cs.AI 62%

Compressing Visual-linguistic Model via Knowledge Distillation

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, Zicheng Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01394 2021-04-06 cs.CV cs.CL cs.LG 62%

MMBERT: Multimodal BERT Pretraining for Improved Medical VQA

Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, CV Jawahar

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11537 2021-03-23 cs.LG cs.CV 62%

How to Design Sample and Computationally Efficient VQA Models

Karan Samel, Zelin Zhao, Binghong Chen, Kuan Wang, Robin Luo, Le Song

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05900 2021-03-11 cs.CV cs.AI 62%

RL-CSDia: Representation Learning of Computer Science Diagrams

Shaowei Wang, LingLing Zhang, Xuan Luo, Yi Yang, Xin Hu, Jun Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.06793 2021-02-16 cs.CV cs.AI cs.CL 62%

Unanswerable Questions about Images and Texts

Ernest Davis

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 4 figures

Journal ref Frontiers in Artificial Intelligence: Language and Computation. July 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.00424 2021-02-02 cs.CL cs.CV cs.LG 62%

An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games

Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted paper for the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14891 2021-01-01 cs.CV cs.LG 62%

Detecting Hate Speech in Multi-modal Memes

Abhishek Das, Japsimar Singh Wahi, Siyao Li

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07214 2020-10-30 cs.LG cs.CL cs.CV stat.ML 62%

Sparse and Continuous Attention Mechanisms

André F. T. Martins, António Farinhas, Marcos Treviso, Vlad Niculae, Pedro M. Q. Aguiar, Mário A. T. Figueiredo

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted for spotlight presentation at NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10802 2020-10-22 cs.CV cs.LG 62%

Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies

Itai Gat, Idan Schwartz, Alexander Schwing, Tamir Hazan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10354 2020-10-13 cs.CV cs.CL cs.LG 62%

Unsupervised Keyword Extraction for Full-sentence VQA

Kohei Uehara, Tatsuya Harada

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments EMNLP 2020 workshop: NLP Beyond Text (NLPBT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00562 2020-10-02 cs.CL cs.AI cs.CV 62%

ISAAQ -- Mastering Textbook Questions with Pre-trained Transformers and Bottom-Up and Top-Down Attention

Jose Manuel Gomez-Perez, Raul Ortega

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted for publication as a long paper in EMNLP2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.00124 2020-08-31 cs.CV cs.CL cs.HC cs.LG 62%

A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop

Jihyeon Lee, Sho Arora

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 9 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07194 2020-08-05 cs.CL cs.CV cs.IR cs.LG cs.MM 62%

Recommending Themes for Ad Creative Design via Visual-Linguistic Representations

Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, Wei Wang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 7 pages, 8 figures, 2 tables, accepted by The Web Conference 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11740 2020-07-21 cs.CV cs.CL cs.LG 62%

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, Jingjing Liu

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.05608 2020-07-14 cs.CV cs.CL cs.LG 62%

Image Captioning with Compositional Neural Module Networks

Junjiao Tian, Jean Oh

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments International Joint Conference on Artificial Intelligence (IJCAI-19)

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13073 2020-07-14 cs.CV cs.CL cs.LG 62%

A Novel Attention-based Aggregation Function to Combine Vision and Language

Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02509 2020-07-14 cs.LG cs.CV cs.NE 62%

REMIND Your Neural Network to Prevent Catastrophic Forgetting

Tyler L. Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, Christopher Kanan

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments To appear in the European Conference on Computer Vision (ECCV-2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02833 2020-07-07 cs.CV cs.LG 62%

Eliminating Catastrophic Interference with Biased Competition

Amelia Elizabeth Pollard, Jonathan L. Shapiro

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00900 2020-07-03 cs.CV cs.AI cs.HC 62%

The Impact of Explanations on AI Competency Prediction in VQA

Kamran Alipour, Arijit Ray, Xiao Lin, Jurgen P. Schulze, Yi Yao, Giedrius T. Burachas

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Submitted to HCCAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16917 2020-07-01 cs.AI cs.LG 62%

Ontology-guided Semantic Composition for Zero-Shot Learning

Jiaoyan Chen, Freddy Lecue, Yuxia Geng, Jeff Z. Pan, Huajun Chen

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

Comments Accepted by KR 2020 - 17th International Conference on Principles of Knowledge Representation and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.07538 2020-06-01 cs.CV cs.CL cs.LG 62%

Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing

Vedika Agarwal, Rakshith Shetty, Mario Fritz

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09241 2020-05-20 cs.CV cs.LG 62%

On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law

Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.09034 2020-04-21 cs.CV cs.LG 62%

Learning What Makes a Difference from Counterfactual Examples and Gradient Supervision

Damien Teney, Ehsan Abbasnedjad, Anton van den Hengel

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08530 2020-02-19 cs.CV cs.CL cs.LG 62%

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, Jifeng Dai

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted by ICLR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03712 2020-01-14 cs.CV cs.CL cs.LG 62%

MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding

Geondo Park, Chihye Han, Wonjun Yoon, Daeshik Kim

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments Accepted by the 2020 IEEE Winter Conference on Applications of Computer Vision (WACV 20), 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04964 2020-01-07 cs.MM cs.CV cs.IR cs.LG 62%

Multi-modal Deep Analysis for Multimedia

Wenwu Zhu, Xin Wang, Hongzhi Li

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 25 pages, 39 figures, IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03063 2019-12-09 cs.CV cs.CL cs.LG cs.NE 62%

Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks

Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.08618 2019-11-21 cs.CV cs.LG cs.MM eess.IV 62%

Explanation vs Attention: A Two-Player Game to Obtain Attention for VQA

Badri N. Patro, Anupriy, Vinay P. Namboodiri

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments AAAI-2020(Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏