arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2602.14983 2026-02-17 cs.LG 82%

Orthogonalized Multimodal Contrastive Learning with Asymmetric Masking for Structured Representations

正交化多模态对比学习与不对称掩码用于结构化表示

Carolin Cissee, Raneen Younis, Zahra Ahmadi

机构 * Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School(图林根工业大学和汉诺威医学院医学信息学研究所) Lower Saxony Center for AI and Causal Methods in Medicine (CAIMed)(下萨克森人工智能与因果医学中心(CAIMed))

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 COrAL通过正交化多模态对比学习与不对称掩码,显式保留冗余、独特和协同信息,提升多模态表示的稳定性与全面性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14107 2026-02-17 cs.DC 82%

ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies

ML-ECS: 一种用于边缘-云协同的协同多模态学习框架

Yuze Liu, Shibo Chu, Tiehua Zhang, Hao Zhou, Zhishu Shen, Jinze Wang, Jianzhong Qi, Feng Xia

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 ML-ECS提出了一种协同多模态学习框架,通过跨模态对比学习、自适应多模态调节、模态感知聚合和SLM增强的CCL,实现边缘-云协同中的高效知识转移与性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09510 2026-02-17 cs.IR 82%

MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval

MRMR:一个面向推理密集型多模态检索的现实且专家级多学科基准

Siyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao, Tingyu Song, Arman Cohan, Anh Tuan Luu, Chen Zhao

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)

AI总结 MRMR是一个面向推理密集型多模态检索的现实且专家级多学科基准,通过多领域查询和混合模态数据推动检索模型的改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09221 2026-02-12 cs.HC cs.LG 82%

CMCRD: Cross-Modal Contrastive Representation Distillation for Emotion Recognition

CMCRD:跨模态对比表示蒸馏用于情绪识别

Siyuan Kan, Huanyu Wu, Zhenyao Cui, Fan Huang, Xiaolong Xu, Dongrui Wu

机构 * Key Laboratory of Image Processing and Intelligent Control, Ministry of Education, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(图像处理与智能控制重点实验室、教育部、自动化学院、华中科技大学) Wuhan Institute of Digital Engineering(武汉数字工程研究所) Shanghai Jiao Tong University(上海交通大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract)

AI总结 CMCRD通过跨模态对比表示蒸馏提升情绪识别准确性,减少多模态数据需求,实验表明在EEG和眼动数据间相互辅助训练可提高分类准确率约6.2%。

Journal ref IEEE Trans. on Affective Computing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04021 2026-02-05 cs.LG q-bio.QM stat.ML 82%

Group Contrastive Learning for Weakly Paired Multimodal Data

用于弱配对多模态数据的组对比学习

Aditya Gorla, Hugues Van Assel, Jan-Christian Huetter, Heming Yao, Kyunghyun Cho, Aviv Regev, Russell Littman

机构 * Research and Early Development (gRED), Genentech(Genentech 研究与早期发展部门) UCLA(加州大学洛杉矶分校) Biology Research | AI Development (BRAID), Genentech(Genentech 生物研究 | 人工智能开发部门) Genentech Computational Sciences NYU(Genentech 计算科学与纽约大学)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);cross-modal(abstract)

AI总结 GROOVE通过组级对比学习方法,有效处理弱配对多模态数据,提升跨模态匹配与填补任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09114 2026-02-03 cs.LG 82%

TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval

TRACE: 在上下文中对时间序列进行建模以实现多模态嵌入与检索

Jialin Chen, Ziyu Zhao, Gaukhar Nurbek, Aosong Feng, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, Rex Ying

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 TRACE通过在上下文中对时间序列进行建模,实现多模态嵌入与检索,提升下游任务的预测精度和可解释性,同时作为强大的独立编码器优化上下文感知表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20028 2026-01-29 cs.LG 82%

Decomposing multimodal embedding spaces with group-sparse autoencoders

用组稀疏自编码器分解多模态嵌入空间

Chiraag Kaushik, Davis Barch, Andrea Fanelli

机构 * School of Electrical and Computer Engineering(电气与计算机工程学院) Georgia Institute of Technology(佐治亚理工学院) Dolby Laboratories(杜比实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出基于组稀疏正则化的多模态嵌入分解方法,通过跨模态随机遮蔽提升多模态对齐,减少死神经元并增强语义性。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10096 2026-01-22 cs.LG 82%

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

多语言到多模态(M2M):通过单语文本解锁新语言

Piyush Singh Pasi

机构 * Amazon(亚马逊)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)

AI总结 M2M通过单语文本学习多语言多模态对齐,实现多语言文本到图像检索的零样本迁移

Comments EACL 2026 Findings accepted. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00725 2025-12-02 cs.LG 82%

ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering

ESMC: 基于多模态大语言模型的可解释多聚类嵌入选择

Xinyue Wang, Yuheng Jia, Hui Liu, Junhui Hou

专题命中 跨模态检索 :MLLM(title,abstract);multi-modal(abstract)

AI总结 ESMC利用多模态大语言模型实现用户驱动的可解释多聚类,通过嵌入选择和伪标签学习提升聚类准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17507 2025-12-02 cs.IR 82%

CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling

CART:基于粗到细语义建模的生成式跨模态检索框架

Minghui Fang, Shengpeng Ji, Jialong Zuo, Hai Huang, Yan Xia, Jieming Zhu, Xize Cheng, Xiaoda Yang, Wenrui Liu, Gang Wang, Zhenhua Dong, Zhou Zhao

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract)

AI总结 CART通过生成式模型和粗到细语义建模提升跨模态检索的效率与性能。

Comments Updated a few baseline metrics based on original publications. Conclusions unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19509 2025-11-26 cs.LG 82%

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

TouchFormer: 一种基于变换器的鲁棒多模态材料感知框架

Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong, Xuexin Liu, Zhuojun Zou, Haoyue Yang, Lin Shu, Jie Hao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 TouchFormer通过模态自适应门控和注意力机制提升多模态材料感知的鲁棒性,并在细粒度分类任务中实现性能提升。

Comments 9 pages, 7 figures, Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09250 2025-11-13 cs.IR 82%

NeuroCLIP: Brain-Inspired Prompt Tuning for EEG-to-Image Multimodal Contrastive Learning

Jiyuan Wang, Li Zhang, Haipeng Lin, Qile Liu, Gan Huang, Ziyu Li, Zhen Liang, Xia Wu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05325 2025-11-10 cs.LG 82%

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

Janet Jenq, Hongda Shen

机构 * PitchBook USA(PitchBook美国公司)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08828 2025-11-10 cs.IR cs.AI cs.CL cs.CV 82%

MMDocIR: Benchmarking Multimodal Retrieval for Long Documents

Kuicai Dong, Yujing Chang, Xin Deik Goh, Dexun Li, Ruiming Tang, Yong Liu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Paper accepted to EMNLP-2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02728 2025-11-07 cs.RO 82%

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

Lingfeng Zhang, Erjia Xiao, Yuchen Zhang, Haoxiang Fu, Ruibin Hu, Yanbiao Ma, Wenbo Ding, Long Chen, Hangjun Ye, Xiaoshuai Hao

机构 * Tsinghua University(清华大学) Xiaomi EV(小米电动车) Georgia Institute of Technology(佐治亚理工学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Renmin University of China(中国人民大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02371 2025-11-05 cs.LG 82%

LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment

Rohan Wandre, Yash Gajewar, Namrata Patel, Vivek Dhalkari

机构 * Dept. of Computer Engineering(计算机工程系) SIES Graduate School of Technology(SIES技术研究生学院) Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12815 2025-11-04 cs.LG cs.AI cs.CL cs.CV 82%

Learning to Steer: Input-dependent Steering for Multimodal LLMs

Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny, Alasdair Newson, Matthieu Cord

机构 * ISIR, Sorbonne Université(ISIR,索邦大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08809 2025-11-04 cs.LG 82%

Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models

Wei Chen, Xin Yan, Bin Wen, Fan Yang, Tingting Gao, Di Zhang, Long Chen

机构 * HKUST(香港科技大学) University of Waterloo(滑铁卢大学) Kuaishou Technology(快手科技)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract)

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15543 2025-10-20 cs.CL cs.AI cs.IR cs.MM 82%

MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval

Qiyu Wu, Shuyang Cui, Satoshi Hayakawa, Wei-Yao Wang, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14592 2025-10-17 cs.LG cs.IR 82%

Multimodal RAG for Unstructured Data:Leveraging Modality-Aware Knowledge Graphs with Hybrid Retrieval

Rashmi R, Vidyadhar Upadhya

机构 * National Institute of Technology Karnataka, Surathkal, India(印度卡纳塔克国家理工学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

Comments 12 pages, 6 figures, submitted for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13856 2025-10-17 cs.CL cs.AI cs.CV 82%

Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA

A H M Rezaul Karim, Ozlem Uzuner

机构 * George Mason University(乔治·马歇尔大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20152 2025-10-02 cs.CV cs.AI cs.CL 82%

MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models

Kai Sun, Yushi Bai, Zhen Yang, Jiajie Zhang, Ji Qi, Lei Hou, Juanzi Li

机构 * Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21339 2025-09-29 cs.IR cs.AI cs.CV cs.MM 82%

Cross-Modal Retrieval with Cauchy-Schwarz Divergence

Jiahao Zhang, Wenzhe Yin, Shujian Yu

机构 * The HongKong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Amsterdam(阿姆斯特丹大学) Vrije Universiteit Amsterdam(自由大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACMMM-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08216 2025-09-11 cs.IR 82%

Vector embedding of multi-modal texts: a tool for discovery?

Beth Plale, Sai Navya Jyesta, Sachith Withana

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24073 2025-08-27 cs.AI cs.CL cs.CV 82%

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation

Chan-Wei Hu, Yueqi Wang, Shuo Xing, Chia-Ju Chen, Suofei Feng, Ryan Rossi, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯农工大学) University of California, Berkeley(加州大学伯克利分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) Stanford University(斯坦福大学) Adobe Research(Adobe研究院)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11673 2025-08-19 cs.LG cs.AI cs.CV cs.MM 82%

Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning

Haojie Zhang, Yixiong Liang, Hulin Kuang, Lihui Cen, Zhe Qu, Yigang Cen, Min Zeng, Shichao Kan

机构 * School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) School of Automation, Central South University(自动化学院,中南大学) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 10 pages, 3 figures, submitted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11141 2025-08-18 cs.CV cs.AI cs.CL 82%

A Cross-Modal Rumor Detection Scheme via Contrastive Learning by Exploring Text and Image internal Correlations

Bin Ma, Yifei Zhang, Yongjin Xian, Qi Li, Linna Zhou, Gongxun Miao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02538 2025-08-05 cs.IR 82%

Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal Retrieval

Zhengxin Pan, Haishuai Wang, Fangyu Wu, Peng Zhang, Jiajun Bu

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18009 2025-07-25 cs.CV cs.AI cs.CL 82%

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

Jake R. Patock, Nicole Catherine Lewis, Kevin McCoy, Christina Gomez, Canling Chen, Lorenzo Luzi

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19329 2025-06-25 cs.LG 82%

Contrastive Cross-Modal Learning for Infusing Chest X-ray Knowledge into ECGs

Vineet Punyamoorty, Aditya Malusare, Vaneet Aggarwal

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏