arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2109.01797 2021-09-07 cs.AI 83%

Hybrid Contrastive Learning of Tri-Modal Representation for Multimodal Sentiment Analysis

Sijie Mai, Ying Zeng, Shuangjia Zheng, Haifeng Hu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.14572 2021-08-10 cs.CV 83%

Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining

Xunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei, Minlong Lu, Yichi Zhang, Hang Xu, Xiaodan Liang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02417 2021-08-06 cs.CV 83%

Structured Multi-modal Feature Embedding and Alignment for Image-Sentence Retrieval

Xuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 9 pages, 7 figures, Accepted by ACM MM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.00724 2021-08-03 cs.CV cs.IR 83%

Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval Service

Zhongwei Xie, Ling Liu, Yanzhao Wu, Lin Li, Luo Zhong

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments accepted by IEEE Transactions on Services Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11149 2021-06-02 cs.CV cs.IR 83%

Compositional Learning of Image-Text Query for Image Retrieval

Muhammad Umer Anwaar, Egor Labintcev, Martin Kleinsteuber

专题命中 跨模态检索 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Published at IEEE WACV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04242 2021-05-11 stat.ML cs.CV cs.LG 83%

T-EMDE: Sketching-based global similarity for cross-modal retrieval

Barbara Rychalska, Mikolaj Wieczorek, Jacek Dabrowski

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 10 pages,5 figures, 4 tables, 1 code snippet

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.13748 2021-04-29 cs.IR cs.MM 83%

QuTI! Quantifying Text-Image Consistency in Multimodal Documents

Matthias Springstein, Eric Müller-Budack, Ralph Ewerth

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments Accepted for publication in: International ACM SIGIR Conference on Research and Development in Information Retrieval 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12836 2021-04-28 cs.CV 83%

Multimodal Contrastive Training for Visual Representation Learning

Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, Baldo Faieta

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04794 2021-04-22 eess.IV cs.CV cs.IR 83%

CMIR-NET : A Deep Learning Based Model For Cross-Modal Retrieval In Remote Sensing

Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05231 2021-03-03 cs.CV 83%

Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders

Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, Stéphane Marchand-Maillet

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted in ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM). arXiv admin note: text overlap with arXiv:2004.09144

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.09144 2021-01-27 cs.CV 83%

Transformer Reasoning Network for Image-Text Matching and Retrieval

Nicola Messina, Fabrizio Falchi, Andrea Esuli, Giuseppe Amato

专题命中 跨模态检索 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Presented at ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.08899 2020-11-19 cs.CV cs.LG 83%

Multimodal Prototypical Networks for Few-shot Learning

Frederik Pahde, Mihai Puscas, Tassilo Klein, Moin Nabi

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments To appear at WACV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.08189 2020-10-19 cs.CV 83%

New Ideas and Trends in Deep Multimodal Content Understanding: A Review

Wei Chen, Weiping Wang, Li Liu, Michael S. Lew

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.01391 2020-01-03 cs.CV 83%

Semi-Supervised Cross-Modal Retrieval with Label Prediction

Devraj Mandal, Pramod Rao, Soma Biswas

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Updated Version of the Paper has been accepted in IEEE Transactions on Multimedia {https://ieeexplore.ieee.org/document/8907496/}

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.13689 2019-10-01 cs.MM 83%

Diachronic Cross-modal Embeddings

David Semedo, João Magalhães

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

Comments To appear in ACM MM 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.01963 2019-09-13 cs.CV 83%

MTFH: A Matrix Tri-Factorization Hashing Framework for Efficient Cross-Modal Retrieval

Xin Liu, Zhikai Hu, Haibin Ling, Yiu-ming Cheung

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 16 pages, accepted by IEEE T-PAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence,2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.07673 2019-08-22 cs.IR cs.MM 83%

Learning Joint Embedding for Cross-Modal Retrieval

Donghuo Zeng

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.MM

Comments 3 pages, 1 figure, Submitted to ICDM2019 Ph.D. Forum session

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.07388 2019-08-21 cs.LG cs.CV 83%

Cross-modal Zero-shot Hashing

Xuanwu Liu, Zhao Li, Jun Wang, Guoxian Yu, Carlotta Domeniconi, Xiangliang Zhang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.04979 2019-08-15 cs.LG cs.CV cs.IR stat.ML 83%

Harmonized Multimodal Learning with Gaussian Process Latent Variable Models

Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04402 2019-07-18 cs.CV 83%

Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval

Yale Song, Mohammad Soleymani

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2019. Includes supplementary material. Have updated results on TGIF and MRW

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.13151 2019-06-13 cs.MM 83%

Semantic Modeling of Textual Relationships in Cross-Modal Retrieval

Jing Yu, Chenghao Yang, Zengchang Qin, Zhuoqian Yang, Yue Hu, Weifeng Zhang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.MM

Comments To appear in KSEM 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.02997 2018-05-09 cs.CV 83%

Category-Based Deep CCA for Fine-Grained Venue Discovery from Multimodal Data

Yi Yu, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.06582 2017-10-19 cs.MM cs.LG stat.ML 83%

Learning Social Image Embedding with Deep Multimodal Attention Networks

Feiran Huang, Xiaoming Zhang, Zhoujun Li, Tao Mei, Yueying He, Zhonghua Zhao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Journal ref Proceedings of Thematic Workshops of the 25th ACM Multimedia 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.09888 2017-05-30 cs.CV 83%

Cross-modal Subspace Learning for Fine-grained Sketch-based Image Retrieval

Peng Xu, Qiyue Yin, Yongye Huang, Yi-Zhe Song, Zhanyu Ma, Liang Wang, Tao Xiang, W. Bastiaan Kleijn, Jun Guo

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.06697 2017-02-21 cs.CV 83%

Correlation Hashing Network for Efficient Cross-Modal Retrieval

Yue Cao, Mingsheng Long, Jianmin Wang, Philip S. Yu

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.06098 2016-12-20 cs.CV 83%

Cross-Modal Manifold Learning for Cross-modal Retrieval

Sailesh Conjeti, Anees Kazi, Nassir Navab, Amin Katouzian

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14698 2024-08-30 cs.IR cs.AI cs.CL cs.CV 83%

Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express

Cherag Aroraa, Tracy Holloway King, Jayant Kumar, Yi Lu, Sanat Sharma, Arvind Srikantan, David Uvalle, Josep Valls-Vargas, Harsha Vardhan

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI;multimodal(comments)

Comments CIKM 2024 (International Conference on Information and Knowledge Management), Multimodal Search and Recommendations Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07032 2026-06-08 cs.CV cs.AI 新提交 82%

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets

前所未见:基于一致视频源数据集的真正零样本组合图像检索基准测试

Zhenyu Yang, Zemin Du, Shengsheng Qian, Changsheng Xu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :MLLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 针对现有零样本组合图像检索数据集存在参考与目标图像不相关、非真正零样本的问题,提出ZeroSight基准,包含来自视频的一致参考-目标对和训练无关的MLLM驱动方法SC4CIR,通过三重对称一致性检查识别难负样本,实验表明现有方法性能被高估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18786 2026-08-17 cs.IR 版本更新 82%

Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation

超越噪声信号:用于多模态序列推荐的双层去噪

Jie Luo, Qi Jin, Xinming Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 研究多模态序列推荐中的双重噪声困境,提出DDMSR框架,从特征拓扑和序列频率角度净化信号,设计基于图的特征去噪与频域序列去噪模块,纳入多模态对比对齐目标,实验证明该框架性能优于基线。

Comments Accepted by ACM MM 2026. 12 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16878 2026-08-14 cs.LG 版本更新 82%

OC-Distill: Ontology-aware Contrastive Learning with Cross-Modal Distillation for ICU Risk Prediction

OC-Distill:基于跨模态蒸馏的面向本体的对比学习用于ICU风险预测

Zhongyuan Liang, Junhyung Jo, Hyang-Jung Lee, Sang Kyu Kim, Irene Y. Chen

机构 * UC Berkeley(伯克利大学) UCSF(旧金山大学) Samsung Advanced Institute of Technology (SAIT)(三星先进技术研究所)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract)

AI总结 OC-Distill通过两阶段框架提升ICU风险预测性能,利用本体感知对比学习和跨模态知识蒸馏,提高模型对患者相似性的捕捉能力及预测效率。

详情

展开后加载摘要…

URL PDF HTML 收藏