arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2308.12320 2023-11-22 cs.CV 83%

Understanding Dark Scenes by Contrasting Multi-Modal Observations

Xiaoyu Dong, Naoto Yokoya

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments WACV2024. Supp: https://drive.google.com/file/d/1Cfn70-Y9JXUuVcFNTk8162w4-YA32W-K/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09696 2023-10-17 cs.AI 83%

Progressive Evidence Refinement for Open-domain Multimodal Retrieval Question Answering

Shuwen Yang, Anran Wu, Xingjiao Wu, Luwei Xiao, Tianlong Ma, Cheng Jin, Liang He

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16141 2023-09-29 cs.CV 83%

Align before Search: Aligning Ads Image to Text for Accurate Cross-Modal Sponsored Search

Yuanmin Tang, Jing Yu, Keke Gai, Yujing Wang, Yue Hu, Gang Xiong, Qi Wu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04961 2023-09-12 cs.IR cs.CV 83%

Multi-modal Extreme Classification

Anshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy, Seba Kuruvilla, Jitendra Ajmera, Keng-hao Chang, Sumeet Agarwal, Purushottam Kar, Manik Varma

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15273 2023-08-30 cs.CV 83%

Cross-Modal Retrieval Meets Inference:Improving Zero-Shot Classification with Cross-Modal Retrieval

Seongha Eom, Namgyu Ho, Jaehoon Oh, Se-Young Yun

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13077 2023-08-28 cs.CV 83%

Preserving Modality Structure Improves Multi-Modal Learning

Swetha Sirnam, Mamshad Nayeem Rizve, Nina Shvetsova, Hilde Kuehne, Mubarak Shah

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07341 2023-07-20 cs.IR cs.CV 83%

PiTL: Cross-modal Retrieval with Weakly-supervised Vision-language Pre-training via Prompting

Zixin Guo, Tzu-Jui Julius Wang, Selen Pehlivan, Abduljalil Radman, Jorma Laaksonen

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Journal ref SIGIR, 2023, 2261-2265

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00316 2023-07-04 cs.LG cs.AI 83%

SHARCS: Shared Concept Space for Explainable Multimodal Learning

Gabriele Dominici, Pietro Barbiero, Lucie Charlotte Magister, Pietro Liò, Nikola Simidjievski

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07284 2023-06-14 cs.CV 83%

Align and Attend: Multimodal Summarization with Dual Contrastive Losses

Bo He, Jun Wang, Jielin Qiu, Trung Bui, Abhinav Shrivastava, Zhaowen Wang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10254 2023-05-01 cs.CV 83%

Image-text Retrieval via Preserving Main Semantics of Vision

Xu Zhang, Xinzheng Niu, Philippe Fournier-Viger, Xudong Dai

专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 6 pages, 3 figures, accepted by ICME2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03391 2023-04-10 cs.CV 83%

Exposing and Mitigating Spurious Correlations for Cross-Modal Retrieval

Jae Myung Kim, A. Sophia Koepke, Cordelia Schmid, Zeynep Akata

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments CVPR'23 MULA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12489 2023-03-23 cs.LG cs.AI cs.CL cs.CV cs.MM 83%

Few-shot Multimodal Multitask Multilingual Learning

Aman Chadha, Vinija Jain

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04267 2023-03-17 cs.CV cs.LG 83%

Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval

Mustafa Shukor, Nicolas Thome, Matthieu Cord

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Code: https://github.com/mshukor/VLPCook

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00312 2023-03-02 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

Multimodal Analogical Reasoning over Knowledge Graphs

Ningyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang, Shumin Deng, Huajun Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ICLR 2023. The project website is https://zjunlp.github.io/project/MKG_Analogy/introduction.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.11429 2023-01-24 cs.CV 83%

A Novel Self-Supervised Cross-Modal Image Retrieval Method In Remote Sensing

Gencer Sumbul, Markus Müller, Begüm Demir

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted at IEEE International Conference on Image Processing (ICIP) 2022. Our code is available at https://git.tu-berlin.de/rsim/SS-CM-RSIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03382 2022-12-20 cs.CV 83%

Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations

Jie Jiang, Shaobo Min, Weijie Kong, Dihong Gong, Hongfa Wang, Zhifeng Li, Wei Liu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11190 2022-11-22 cs.CV 83%

Cross-Modal Contrastive Learning for Robust Reasoning in VQA

Qi Zheng, Chaoyue Wang, Daqing Liu, Dadong Wang, Dacheng Tao

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05232 2022-11-11 cs.CV cs.LG 83%

MuMIC -- Multimodal Embedding for Multi-label Image Classification with Tempered Sigmoid

Fengjun Wang, Sarai Mizrachi, Moran Beladev, Guy Nadav, Gil Amsalem, Karen Lastmann Assaraf, Hadas Harush Boker

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02053 2022-10-21 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, James Zou

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published at NeurIPS 2022. Code and data are available at https://modalitygap.readthedocs.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15270 2022-10-03 cs.CV 83%

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Bin Shan, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05123 2022-08-09 cs.AI 83%

ICAF: Iterative Contrastive Alignment Framework for Multimodal Abstractive Summarization

Zijian Zhang, Chang Shu, Youxin Chen, Jing Xiao, Qian Zhang, Lu Zheng

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted by WCCI-IJCNN 2022 as an oral paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15537 2022-07-01 eess.AS cs.SD 83%

On Metric Learning for Audio-Text Cross-Modal Retrieval

Xinhao Mei, Xubo Liu, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 eess.AS

Comments 5 pages, accepted to InterSpeech2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.11047 2022-04-18 cs.CV 83%

Cross-Modal Coherence for Text-to-Image Retrieval

Malihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia, Vladimir Pavlovic, Matthew Stone

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments This paper is published in AAAI-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00325 2022-04-05 cs.CV 83%

CAT-Det: Contrastively Augmented Transformer for Multi-modal 3D Object Detection

Yanan Zhang, Jiaxin Chen, Di Huang

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16778 2022-04-01 cs.CV 83%

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Kun Yao, Jie Chen, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13815 2022-03-28 cs.CV 83%

Versatile Multi-Modal Pre-Training for Human-Centric Perception

Fangzhou Hong, Liang Pan, Zhongang Cai, Ziwei Liu

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments CVPR 2022; Project Page https://hongfz16.github.io/projects/HCMoCo.html; Codes available at https://github.com/hongfz16/HCMoCo

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08125 2022-01-21 cs.CV 83%

Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing

Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07209 2021-12-15 cs.IR cs.AI cs.LG 83%

ACE-BERT: Adversarial Cross-modal Enhanced BERT for E-commerce Retrieval

Boxuan Zhang, Chao Wei, Yan Jin, Weiru Zhang

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07180 2021-11-16 cs.CL cs.LG 83%

Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning

Yizhen Zhang, Minkyu Choi, Kuan Han, Zhongming Liu

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments 10 pages, 7 figures, 1 appendix, to be published in Neurips 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01345 2021-10-01 cs.CV cs.IR cs.LG 83%

Cross-Modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Representation Learning

Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏