arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3433 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3433 篇

2407.21439 2024-09-26 cs.AI cs.CL cs.LG 88%

MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training

Zhanpeng Chen, Chengjin Xu, Yiyan Qi, Jian Guo

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14556 2024-08-15 cs.CL cs.AI 88%

Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis

Ronghao Lin, Haifeng Hu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06355 2024-03-12 cs.CL cs.CV 88%

Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment

Ming Zhang, Ke Chang, Yunfang Wu

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 4 figures, accepted by LREC-COLING 2024(main conference, long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10071 2022-10-07 cs.CV cs.AI 88%

Contrastive Learning with Cross-Modal Knowledge Mining for Multimodal Human Activity Recognition

Razvan Brinzea, Bulat Khaertdinov, Stylianos Asteriadis

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments to be published in IEEE WCCI 2022 (IJCNN 2022 track)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07793 2018-08-24 cs.MM cs.CV cs.IR 88%

Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval

Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis, Amit K. Roy-Chowdhury

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(title,abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11043 2026-05-19 cs.AI 88%

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models

EmergentBridge: 提升统一多模态嵌入模型中的零样本跨模态迁移

Jincheng Xie, Xingchen Xiao, Runheng Liu, Zhongyi Huang, Yu Zheng, Heyan Huang

机构 * Tsinghua University(清华大学) School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) JD iCity, JD Technology, JD Intelligent Cities Research(京东i城、京东科技、京东智能城市研究院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出EmergentBridge框架,通过学习噪声桥梁锚点和子空间对齐,提升未配对模态对的零样本迁移性能,无需 exhaustive pairwise 监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01946 2026-03-24 cs.LG cond-mat.mtrl-sci cs.AI physics.chem-ph 88%

COFAP: A Universal Framework for COFs Adsorption Prediction through Designed Multi-Modal Extraction and Cross-Modal Synergy

COFAP:通过设计的多模态提取和跨模态协同的通用COFs吸附预测框架

Zihan Li, Mingyang Wan, Mingyu Gao, Xishi Tai, Zhongshan Chen, Xiangke Wang, Feifan Zhang

机构 * College of Science, College of Information and Electrical Engineering(科学学院,信息与电气工程学院) China Agricultural University(中国农业大学) Qingdao Institute of Software, College of Computer Science and Technology(软件研究所,计算机科学与技术学院) China University of Petroleum (East China)(中国石油大学(华东)) Weifang university(潍坊大学) College of Environmental Science and Engineering(环境科学与工程学院) North China Electric Power University(华北电力大学) College of Science(科学学院)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出COFAP框架,通过深度学习提取多模态结构和化学特征,并利用跨模态注意力机制融合特征,实现高效COFs吸附预测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11151 2026-01-19 cs.IR cs.AI 88%

Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation

跨模态注意力网络与双图学习在多模态推荐中的应用

Ji Dai, Quan Fang, Jun Hu, Desheng Cai, Yang Yang, Can Zhao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学) Tianjin University of Technology(天津工业大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管重大科技专项实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 CRANE通过双图学习和递归注意力机制,解决多模态推荐中的浅层融合和不对称特征处理问题,提升推荐性能。

Comments Accepted to ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19257 2025-11-25 cs.CR cs.AI cs.LG 88%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02745 2025-10-28 cs.CV 88%

Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

Lanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu, Shiqi Wang

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯) Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12600 2025-09-17 cs.LG cs.AI q-bio.QM 88%

A Multimodal Foundation Model to Enhance Generalizability and Data Efficiency for Pan-cancer Prognosis Prediction

Huajun Zhou, Fengtao Zhou, Jiabo Ma, Yingxue Xu, Xi Wang, Xiuming Zhang, Li Liang, Zhenhui Li, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Hong Kong University of Science and Technology(香港科学与技术大学) Department of Pathology(病理学系) School of Medicine(医学院) Zhejiang University(浙江大学) Nanfang Hospital and School of Basic Medical Sciences(南方医科大学基础医学系) Southern Medical University(南方医学院) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理重点实验室) Jinfeng Laboratory(金凤实验室) Department of Radiology(放射科) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders(神经系统疾病国家重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17079 2025-08-26 cs.IR cs.AI 88%

Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation

Yejin Choi, Jaewoo Park, Janghan Yoon, Saejin Kim, Jaehyun Jeon, Youngjae Yu

机构 * Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07666 2025-08-12 cs.MM 88%

Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts

Xianbing Zhao, Shengzun Yang, Buzhou Tang, Ronghuan Jiang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14348 2025-07-29 cs.CV 88%

Manipulating Multimodal Agents via Cross-Modal Prompt Injection

Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北洋大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Henan University of Science and Technology(河南科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00211 2025-06-09 cs.RO cs.AI cs.LG cs.SY eess.SY 88%

SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models

Jiawei Zhang, Xuan Yang, Taiqi Wang, Yu Yao, Aleksandr Petiushko, Bo Li

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08648 2024-07-12 cs.CV 88%

CAR-MFL: Cross-Modal Augmentation by Retrieval for Multimodal Federated Learning with Missing Modalities

Pranav Poudel, Prashant Shrestha, Sanskar Amgain, Yash Raj Shrestha, Prashnna Gyawali, Binod Bhattarai

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments Accepted at MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09597 2022-10-05 cs.CV 88%

More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching

Yuxiao Chen, Jianbo Yuan, Long Zhao, Tianlang Chen, Rui Luo, Larry Davis, Dimitris N. Metaxas

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09730 2022-04-22 cs.CV 88%

Transformer Decoders with MultiModal Regularization for Cross-Modal Food Retrieval

Mustafa Shukor, Guillaume Couairon, Asya Grechka, Matthieu Cord

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2022, MULA Workshop. Code is available at https://github.com/mshukor/TFood

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08451 2021-11-17 cs.LG cs.AI 88%

Which is Making the Contribution: Modulating Unimodal and Cross-modal Dynamics for Multimodal Sentiment Analysis

Ying Zeng, Sijie Mai, Haifeng Hu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments Updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.05631 2021-05-13 cs.LG cs.CV eess.IV 88%

Cross-Modal and Multimodal Data Analysis Based on Functional Mapping of Spectral Descriptors and Manifold Regularization

Maysam Behmanesh, Peyman Adibi, Jocelyn Chanussot, Sayyed Mohammad Saeed Ehsani

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.03737 2021-05-06 cs.MM cs.IR 88%

Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval

Donghuo Zeng, Yi Yu, Keizo Oyama

专题命中 跨模态检索 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.MM

Comments 21 pages,11 figures

Journal ref ACM TOMM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08064 2026-06-09 cs.MM cs.CV 版本更新 88%

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

PUMA: 基于层剪枝的语言模型,用于具有模态自适应学习的高效统一多模态检索

Yibo Lyu, Rui Shao, Gongwei Chen, Yijie Zhu, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.MM

AI总结 提出PUMA,通过层剪枝自蒸馏减少MLLM参数,并设计模态自适应对比学习损失(MAC-Loss)提升检索效率,在降低资源消耗的同时保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18112 2026-04-30 cs.CL cs.MM 88%

Retrieval-Augmented Multimodal Model for Fake News Detection

增强检索的多模态模型用于虚假新闻检测

Yiheng Li, Weihai Lu, Hanyi Yu, Yue Wang

机构 * University of International Business and Economics(国际商务经济大学) Peking University(北京大学) University of Southern California(南加州大学) Upstart Holdings, Inc.(Upstart Holdings公司)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CL、cs.MM

AI总结 本文提出RAMM模型,通过多模态大语言模型和抽象叙述对齐模块,解决虚假新闻检测中跨实例叙述一致性缺失和领域知识不足的问题,实验验证了其有效性。

Comments Accepted to SIGIR 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07247 2022-02-16 cs.CV cs.AI cs.CL cs.MM cs.SI 88%

CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval

Licheng Yu, Jun Chen, Animesh Sinha, Mengjiao MJ Wang, Hugo Chen, Tamara L. Berg, Ning Zhang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 10 pages, 7 figures. Commerce Multimodal Model towards Real Applications at Facebook

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17359 2025-11-04 cs.IR 88%

MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval

Tianyuan Li, Lei Wang, Ahtamjan Ahmat, Yating Yang, Bo Ma, Rui Dong, Bangju Han

专题命中 跨模态检索 :cross-modal(title,abstract);MLLM(title);multimodal(abstract)

Comments We plan to revise the methodology and update the experimental analysis before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01960 2025-09-23 cs.LG 88%

MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving

Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学) School of Computer Science(计算机学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract)

Comments 17 pages, 13 figures, the second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03450 2025-05-23 cs.LG 88%

MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents

Junpeng Yue, Xinrun Xu, Börje F. Karlsson, Zongqing Lu

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract)

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13898 2025-03-07 cs.LG 88%

Cross-Modal Prototype based Multimodal Federated Learning under Severely Missing Modality

Huy Q. Le, Chu Myaet Thwal, Yu Qiao, Ye Lin Tun, Minh N. H. Nguyen, Eui-Nam Huh, Choong Seon Hong

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract)

Comments 14 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28350 2026-06-30 cs.IR cs.CV 87%

UniCA: Bi-directional Cross-Attention with Positive Similarity Loss for Robust Multi-Modal Retrieval

UniCA:具有正相似性损失的双向交叉注意力用于鲁棒多模态检索

Yini Huang, Wenlong Zhang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Southern Medical University(南方医科大学)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);image-text(abstract)

AI总结 提出UniCA模型,通过双向交叉注意力块和正相似性损失增强跨模态语义对齐,在WebQA基准上混合任务Recall@5提升4.09%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14405 2026-06-02 cs.LG cs.AI 87%

ES-Merging: Biological MLLM Merging via Embedding Space Signals

ES-Merging: 通过嵌入空间信号进行生物多模态大模型合并

Wonbin Lee, Dongki Kim, Sung Ju Hwang

机构 * KAIST(韩国科学技术院) DeepAuto.ai

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出ES-Merging框架,利用嵌入空间信号估计合并系数,实现生物多模态大模型的高效合并,提升跨模态推理和单模态知识保留能力。

详情

展开后加载摘要…

URL PDF HTML 收藏