arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3457 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3457 篇

2109.03084 2021-09-08 cs.CL cs.AI 62%

Learning grounded word meaning representations on similarity graphs

Mariella Dimiccoli, Herwig Wendt, Pau Batlle

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2021 (long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.11950 2021-08-27 cs.CV cs.CL 62%

LocTex: Learning Data-Efficient Visual Representations from Localized Textual Supervision

Zhijian Liu, Simon Stent, Jie Li, John Gideon, Song Han

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments ICCV 2021. Project page: https://loctex.mit.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15049 2021-08-19 cs.CV cs.AI 62%

HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval

Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, Zhongyuan Wang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04533 2021-08-12 cs.CV cs.AI cs.LG 62%

ASMR: Learning Attribute-Based Person Search with Adaptive Semantic Margin Regularizer

Boseung Jeong, Jicheol Park, Suha Kwak

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2021 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09141 2021-06-18 cs.CL cs.CV 62%

Probing Image-Language Transformers for Verb Understanding

Lisa Anne Hendricks, Aida Nematzadeh

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01894 2021-06-16 cs.CL cs.CV cs.IR cs.LG 62%

Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval

Ramon Sanabria, Austin Waters, Jason Baldridge

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01832 2021-04-06 cs.CV cs.AI 62%

Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning

Chaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun, Houqiang Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08673 2021-04-01 cs.CV cs.CL 62%

A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

Linjie Li, Zhe Gan, Jingjing Liu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04863 2021-03-09 cs.CV cs.AI 62%

From Hand-Perspective Visual Information to Grasp Type Probabilities: Deep Learning via Ranking Labels

Mo Han, Sezen Ya{ğ}mur Günay, İlkay Yıldız, Paolo Bonato, Cagdas D. Onal, Taşkın Padır, Gunar Schirner, Deniz Erdo{ğ}muş

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09375 2021-02-19 cs.CV cs.IR cs.MM 62%

Hierarchical Similarity Learning for Language-based Product Image Retrieval

Zhe Ma, Fenghao Liu, Jianfeng Dong, Xiaoye Qu, Yuan He, Shouling Ji

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICASSP 2021. Code and data will be available at https://github.com/liufh1/hsl

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12350 2021-01-27 cs.CL cs.AI 62%

ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces

Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby Lee, Jindong Chen, Blaise Agüera y Arcas

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI Conference on Artificial Intelligence (AAAI-21)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08148 2020-12-16 cs.CL cs.AI 62%

A Response Retrieval Approach for Dialogue Using a Multi-Attentive Transformer

Matteo A. Senese, Alberto Benincasa, Barbara Caputo, Giuseppe Rizzo

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.07000 2020-12-15 cs.AI cs.CL 62%

KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning

Dandan Song, Siyi Ma, Zhanchen Sun, Sicheng Yang, Lejian Liao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05107 2020-12-10 cs.CL cs.CV cs.LG 62%

Towards Zero-shot Cross-lingual Image Retrieval

Pranav Aggarwal, Ajinkya Kale

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02467 2020-10-07 cs.CV cs.CL 62%

Learning Visual-Semantic Embeddings for Reporting Abnormal Findings on Chest X-rays

Jianmo Ni, Chun-Nan Hsu, Amilcare Gentili, Julian McAuley

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments 7 pages, 2 figures, to be published in Findings of EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.12212 2020-09-24 cs.CV cs.CL cs.IR cs.LG 62%

ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual Descriptions

Anurag Roy, Vinay Kumar Verma, Kripabandhu Ghosh, Saptarshi Ghosh

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted in CIKM-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00392 2020-03-03 cs.CV cs.AI 62%

Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

Shizhe Chen, Yida Zhao, Qin Jin, Qi Wu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments To be appeared in CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.09461 2020-02-24 cs.CV cs.MM 62%

Fine-Grained Instance-Level Sketch-Based Video Retrieval

Peng Xu, Kun Liu, Tao Xiang, Timothy M. Hospedales, Zhanyu Ma, Jun Guo, Yi-Zhe Song

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03712 2020-01-14 cs.CV cs.CL cs.LG 62%

MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding

Geondo Park, Chihye Han, Wonjun Yoon, Daeshik Kim

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted by the 2020 IEEE Winter Conference on Applications of Computer Vision (WACV 20), 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10531 2019-11-26 cs.CV cs.MM eess.IV 62%

A Proposal-based Approach for Activity Image-to-Video Retrieval

Ruicong Xu, Li Niu, Jianfu Zhang, Liqing Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments The Thirty-Fourth AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10097 2019-11-25 cs.LG cs.CL cs.CV 62%

HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs

Fangyu Liu, Rongtian Ye, Xun Wang, Shuaipeng Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI-20 (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05978 2019-11-15 cs.CV cs.CL cs.LG 62%

HUSE: Hierarchical Universal Semantic Embeddings

Pradyumna Narayana, Aniket Pednekar, Abishek Krishnamoorthy, Kazoo Sone, Sugato Basu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11119 2019-10-25 cs.CV cs.CL cs.LG 62%

Designovel's system description for Fashion-IQ challenge 2019

Jianri Li, Jae-whan Lee, Woo-sang Song, Ki-young Shin, Byung-hyun Go

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01205 2019-06-05 cs.LG cs.CL cs.CV 62%

A Strong and Robust Baseline for Text-Image Matching

Fangyu Liu, Rongtian Ye

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments 6 pages (excluding references); 2019 ACL Student Research Workshop (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05521 2019-04-30 cs.CV cs.CL cs.LG 62%

UniVSE: Robust Visual Semantic Embeddings via Structured Semantic Representations

Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, Wei-Ying Ma

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments v1 is the full version which is accepted by CVPR 2019. v2 is the short version accepted by NAACL 2019 SpLU-RoboNLP workshop (in non-archival proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.09362 2018-11-27 cs.CL cs.AI 62%

Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors

Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10348 2018-06-28 cs.CL cs.CV 62%

Learning Visually-Grounded Semantics from Contrastive Adversarial Samples

Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, Jian Sun

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

Comments To Appear at COLING 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.08495 2018-03-23 cs.CV cs.AI cs.GR cs.LG 62%

Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings

Kevin Chen, Christopher B. Choy, Manolis Savva, Angel X. Chang, Thomas Funkhouser, Silvio Savarese

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07199 2017-12-21 cs.DB cs.AI cs.CL cs.NE 62%

Cognitive Database: A Step towards Endowing Relational Databases with Artificial Intelligence Capabilities

Rajesh Bordawekar, Bortik Bandyopadhyay, Oded Shmueli

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05851 2017-08-22 cs.CV cs.IR cs.MM 62%

Image2song: Song Retrieval via Bridging Image Content and Lyric Words

Xuelong Li, Di Hu, Xiaoqiang Lu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 13 figures, accepted by ICCV 2017

详情

展开后加载摘要…

URL PDF HTML 收藏