arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2110.12765 2021-10-26 cs.CL cs.AI 62%

"So You Think You're Funny?": Rating the Humour Quotient in Standup Comedy

Anirudh Mittal, Pranav Jeevan, Prerak Gandhi, Diptesh Kanojia, Pushpak Bhattacharyya

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2021 Main Conference (short papers); 4 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06100 2021-10-13 cs.SD cs.MM eess.AS 62%

Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information

Zhongjie Ye, Helin Wang, Dongchao Yang, Yuexian Zou

专题命中 图文多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments 5 pages, 1 figure, accepted by DCASE 2021 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10649 2021-09-23 cs.CV cs.AI 62%

Caption Enriched Samples for Improving Hateful Memes Detection

Efrat Blaier, Itzik Malkiel, Lior Wolf

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12560 2021-08-31 cs.CL cs.CV 62%

QACE: Asking Questions to Evaluate an Image Caption

Hwanhee Lee, Thomas Scialom, Seunghyun Yoon, Franck Dernoncourt, Kyomin Jung

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2021 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02059 2021-08-05 cs.CV cs.MM 62%

Question-controlled Text-aware Image Captioning

Anwen Hu, Shizhe Chen, Qin Jin

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 8 figures, to appear in ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13083 2021-07-29 cs.CV cs.AI 62%

Is Object Detection Necessary for Human-Object Interaction Recognition?

Ying Jin, Yinpeng Chen, Lijuan Wang, Jianfeng Wang, Pei Yu, Zicheng Liu, Jenq-Neng Hwang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.11509 2021-07-27 cs.CV cs.AI 62%

Cycled Compositional Learning between Images and Text

Jongseok Kim, Youngjae Yu, Seunghwan Lee, GunheeKim

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Fashion IQ 2020 challenge winner. Workshop tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03751 2021-07-09 cs.CV cs.AI 62%

Exploiting the relationship between visual and textual features in social networks for image classification with zero-shot deep learning

Luis Lucas, David Tomas, Jose Garcia-Rodriguez

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05918 2021-06-14 cs.CV cs.CL cs.LG 62%

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, Tom Duerig

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments ICML 2021

Journal ref International Conference on Machine Learning 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08560 2021-04-20 cs.CL cs.CV 62%

Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments

Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, Bryan A. Plummer

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted at the workshop on Visually Grounded Interaction and Language (ViGIL) at NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05614 2021-02-15 cs.CV cs.CL 62%

Delving Deeper into the Decoder for Video Captioning

Haoran Chen, Jianmin Li, Xiaolin Hu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 8 pages, 3 figures, European Conference on Artificial Intelligence. ECAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.12404 2020-12-08 cs.CL cs.CV 62%

Visually Grounded Compound PCFGs

Yanpeng Zhao, Ivan Titov

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to EMNLP 2020. Our code is available at https://github.com/zhaoyanpeng/vpcfg

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14901 2020-12-01 cs.CL cs.CV cs.LG cs.NE 62%

Language-Driven Region Pointer Advancement for Controllable Image Captioning

Annika Lindh, Robert J. Ross, John D. Kelleher

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to COLING 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.16363 2020-11-02 cs.CL cs.CV 62%

Domain-Specific Lexical Grounding in Noisy Visual-Textual Documents

Gregory Yauney, Jack Hessel, David Mimno

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Journal ref Published in EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.04295 2020-10-12 cs.LG cs.AI cs.CL cs.HC 62%

Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements

Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, Zhiwei Guan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 16 pages, EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.01725 2020-10-06 cs.CV cs.AI 62%

Attention Guided Semantic Relationship Parsing for Visual Question Answering

Moshiur Farazi, Salman Khan, Nick Barnes

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.04660 2020-07-10 cs.SD cs.LG cs.MM eess.AS 62%

Multi-task Regularization Based on Infrequent Classes for Audio Captioning

Emre Çakır, Konstantinos Drossos, Tuomas Virtanen

专题命中 图文多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11807 2020-06-23 cs.CV cs.CL 62%

Improving Image Captioning with Better Use of Captions

Zhan Shi, Xu Zhou, Xipeng Qiu, Xiaodan Zhu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ACL 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08070 2020-06-16 cs.CV cs.CL 62%

Transform and Tell: Entity-Aware News Image Captioning

Alasdair Tran, Alexander Mathews, Lexing Xie

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Published in CVPR 2020. Code is available at https://github.com/alasdairtran/transform-and-tell and demo is available at https://transform-and-tell.ml

Journal ref The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 13035-13045

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08226 2020-03-24 cs.CV cs.CL 62%

Meshed-Memory Transformer for Image Captioning

Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita Cucchiara

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.02391 2019-12-04 cs.CV cs.AI cs.LG 62%

Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization

Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, Dhruv Batra

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This version was published in International Journal of Computer Vision (IJCV) in 2019; A previous version of the paper was published at International Conference on Computer Vision (ICCV'17)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.02489 2019-09-06 cs.CV cs.CL 62%

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation

Wei Wei, Ling Cheng, Xianling Mao, Guangyou Zhou, Feida Zhu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.08876 2019-06-24 cs.CL cs.CV 62%

Informative Image Captioning with External Sources of Information

Sanqiang Zhao, Piyush Sharma, Tomer Levinboim, Radu Soricut

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.05226 2019-06-13 cs.CL cs.CV cs.LG 62%

Continual and Multi-Task Architecture Search

Ramakanth Pasunuru, Mohit Bansal

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ACL 2019 (12 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.08126 2018-12-20 cs.CV cs.CL cs.LG 62%

Generating Diverse and Meaningful Captions

Annika Lindh, Robert J. Ross, Abhijit Mahalunkar, Giancarlo Salton, John D. Kelleher

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted for presentation at The 27th International Conference on Artificial Neural Networks (ICANN 2018)

Journal ref Artificial Neural Networks and Machine Learning - ICANN 2018 (pp. 176-187). Springer International Publishing

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.08697 2018-09-25 cs.CL cs.CV 62%

Textually Enriched Neural Module Networks for Visual Question Answering

Khyathi Raghavi Chandu, Mary Arpita Pyreddy, Matthieu Felix, Narendra Nath Joshi

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.09137 2018-03-15 cs.NE cs.CL cs.CV 62%

Where to put the Image in an Image Caption Generator

Marc Tanti, Albert Gatt, Kenneth P. Camilleri

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted in JNLE Special Issue: Language for Images (24.3) (expanded with content that was removed from journal paper in order to reduce number of pages), 28 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.00930 2017-08-15 cs.CV cs.AI cs.LG 62%

Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner

Tseng-Hung Chen, Yuan-Hong Liao, Ching-Yao Chuang, Wan-Ting Hsu, Jianlong Fu, Min Sun

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.07770 2017-07-24 cs.CV cs.CL 62%

Captioning Images with Diverse Objects

Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond Mooney, Trevor Darrell, Kate Saenko

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments CVPR 2017 Camera ready version. 17 pages (8 + 9 supplement), 12 figures, 8 tables. Includes project page http://vsubhashini.github.io/noc.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.06676 2017-06-06 cs.CV cs.CL 62%

I2T2I: Learning Text to Image Synthesis with Textual Data Augmentation

Hao Dong, Jingqing Zhang, Douglas McIlwraith, Yike Guo

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments International Conference on Image Processing (ICIP) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏