arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2310.19795 2023-10-31 cs.CV cs.AI cs.LG 84%

SimMMDG: A Simple and Effective Framework for Multi-modal Domain Generalization

Hao Dong, Ismail Nejjar, Han Sun, Eleni Chatzi, Olga Fink

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07362 2023-10-20 cs.CV cs.CL 84%

A scoping review on multimodal deep learning in biomedical images and texts

Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie, Fei Wang, Zhiyong Lu, Yifan Peng

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments This paper has been accepted by the Journal of Biomedical Informatics

Journal ref Journal of Biomedical Informatics, Volume 146, October 2023, 104482

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04456 2023-10-10 cs.CL cs.SD eess.AS 84%

Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in Conversation

Shihao Zou, Xianying Huang, Xudong Shen

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14580 2023-09-27 cs.LG cs.AI cs.CV 84%

CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive Loss

Rakshith Sharma Srinivasa, Jaejin Cho, Chouchang Yang, Yashas Malur Saidutta, Ching-Hua Lee, Yilin Shen, Hongxia Jin

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to Neural Information Processing Systems (NeurIPS) 2023 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05451 2023-09-12 cs.CV cs.MM 84%

Dual-view Curricular Optimal Transport for Cross-lingual Cross-modal Retrieval

Yabing Wang, Shuhui Wang, Hao Luo, Jianfeng Dong, Fan Wang, Meng Han, Xun Wang, Meng Wang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14009 2023-08-29 cs.CV cs.AI 84%

Towards Fast and Accurate Image-Text Retrieval with Self-Supervised Fine-Grained Alignment

Jiamin Zhuang, Jing Yu, Yang Ding, Xiangyan Qu, Yue Hu

专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE Transactions on Multimedia (TMM)

Journal ref IEEE Transactions on Multimedia ( Early Access ), 29 May 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03323 2023-07-18 cs.CV cs.AI cs.CR cs.LG 84%

CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive Learning

Hritik Bansal, Nishad Singhi, Yu Yang, Fan Yin, Aditya Grover, Kai-Wei Chang

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 22 pages. Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10140 2023-05-29 cs.CL cs.CV 84%

Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation

Matthieu Futeral, Cordelia Schmid, Ivan Laptev, Benoît Sagot, Rachel Bawden

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09299 2023-05-17 cs.CV cs.CL 84%

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

Heqing Zou, Meng Shen, Chen Chen, Yuchen Hu, Deepu Rajan, Eng Siong Chng

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05221 2023-04-04 cs.CV cs.AI 84%

REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory

Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, Alireza Fathi

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Published on CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00534 2023-03-02 cs.CV cs.CL 84%

RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training

Zheng Yuan, Qiao Jin, Chuanqi Tan, Zhengyun Zhao, Hongyi Yuan, Fei Huang, Songfang Huang

专题命中 跨模态检索 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01328 2023-02-17 cs.CV cs.AI 84%

DUET: Cross-modal Semantic Grounding for Contrastive Zero-shot Learning

Zhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng, Wen Zhang, Yin Fang, Jeff Z. Pan, Huajun Chen

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2023 (Oral). Repository: https://github.com/zjukg/DUET

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02969 2023-01-12 cs.CV cs.AI 84%

Multi-scale multi-modal micro-expression recognition algorithm based on transformer

Fengping Wang, Jie Li, Chun Qi, Lin Wang, Pan Wang

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05752 2022-12-13 cs.CV cs.AI 84%

Scale-Semantic Joint Decoupling Network for Image-text Retrieval in Remote Sensing

Chengyu Zheng, Ning song, Ruoyu Zhang, Lei Huang, Zhiqiang Wei, Jie Nie

专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06209 2022-09-15 cs.CV cs.IR cs.MM 84%

Look Before You Leap: Improving Text-based Person Retrieval by Learning A Consistent Cross-modal Common Manifold

Zijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan, Chao Liu, Tian Wang, Yifeng Li

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted on ACM MM '22. arXiv admin note: text overlap with arXiv:2209.05773

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12526 2022-08-29 cs.CV cs.MM 84%

Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning

Yabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang, Rui Cai, Xun Wang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2022. Code and data are available at https://github.com/HuiGuanLab/nrccr

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10457 2022-08-22 cs.CV cs.AI 84%

Language Guided Networks for Cross-modal Moment Retrieval

Kun Liu, Huadong Ma, Chuang Gan

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05024 2022-07-14 cs.CV cs.MM 84%

Intra-Modal Constraint Loss For Image-Text Retrieval

Jianan Chen, Lu Zhang, Qiong Wang, Cong Bai, Kidiyo Kpalma

专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04211 2022-07-12 cs.AI cs.CV cs.IR cs.LG eess.IV 84%

BOSS: Bottom-up Cross-modal Semantic Composition with Hybrid Counterfactual Training for Robust Content-based Image Retrieval

Wenqiao Zhang, Jiannan Guo, Mengze Li, Haochen Shi, Shengyu Zhang, Juncheng Li, Siliang Tang, Yueting Zhuang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00733 2022-07-11 cs.CV cs.AI 84%

Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation Learning and Retrieval

Keyu Wen, Zhenshan Tan, Qingrong Cheng, Cheng Chen, Xiaodong Gu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08842 2022-06-20 cs.MM cs.CV cs.DB cs.IR 84%

Entity-Graph Enhanced Cross-Modal Pretraining for Instance-level Product Retrieval

Xiao Dong, Xunlin Zhan, Yunchao Wei, Xiaoyong Wei, Yaowei Wang, Minlong Lu, Xiaochun Cao, Xiaodan Liang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07441 2022-05-23 cs.CV cs.CL cs.IR 84%

COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao, Zhiwu Lu, Ji-Rong Wen

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09268 2022-04-21 cs.LG cs.CL cs.CV cs.IR 84%

Uncertainty-based Cross-Modal Retrieval with Probabilistic Representations

Leila Pishdad, Ran Zhang, Konstantinos G. Derpanis, Allan Jepson, Afsaneh Fazly

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09138 2022-03-18 cs.CV cs.MM 84%

MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering

Yang Ding, Jing Yu, Bang Liu, Yue Hu, Mingxin Cui, Qi Wu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11920 2022-02-22 cs.CV cs.CL 84%

Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval

Gregor Geigle, Jonas Pfeiffer, Nils Reimers, Ivan Vulić, Iryna Gurevych

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments TACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.11843 2022-01-31 cs.CV cs.AI cs.LG 84%

Discriminative Supervised Subspace Learning for Cross-modal Retrieval

Haoming Zhang, Xiao-Jun Wu, Tianyang Xu, Donglin Zhang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.05523 2021-09-14 cs.CV cs.CL 84%

Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text Retrieval

Zhihao Fan, Zhongyu Wei, Zejun Li, Siyuan Wang, Haijun Shan, Xuanjing Huang, Jianqing Fan

专题命中 跨模态检索 :image-text(title);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.16030 2020-11-02 cs.IR cs.MM cs.SD eess.AS 84%

Multimodal Metric Learning for Tag-based Music Retrieval

Minz Won, Sergio Oramas, Oriol Nieto, Fabien Gouyon, Xavier Serra

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM、eess.AS

Comments 5 pages, 2 figures, submitted to ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09641 2020-10-20 cs.MM cs.AI 84%

DIME: An Online Tool for the Visual Comparison of Cross-Modal Retrieval Models

Tony Zhao, Jaeyoung Choi, Gerald Friedland

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.02497 2019-07-30 cs.IR cs.CV cs.MM 84%

Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos

Zhu Zhang, Zhijie Lin, Zhou Zhao, Zhenxin Xiao

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by SIGIR 2019 as a full paper

Journal ref SIGIR, 2019, pages 655-664

详情

展开后加载摘要…

URL PDF HTML 收藏