arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3433 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3433 篇

2602.01753 2026-06-02 cs.CV 87%

ObjEmbed: Towards Universal Multimodal Object Embeddings

ObjEmbed:迈向通用多模态对象嵌入

Shenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu, Xiaohua Xie, Wei-Shi Zheng

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);image-text(abstract);分类 cs.CV

AI总结 提出ObjEmbed模型,通过分解图像为多个区域嵌入(每个对应一个对象)并生成语义和IoU两种互补嵌入,实现细粒度视觉-语言对齐,在视觉定位、局部和全局图像检索等任务中表现优异。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19115 2026-05-12 cs.CV 87%

Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?

生成巨人,检索弱者:为何多模态大语言模型在多模态检索中表现不佳?

Hengyi Feng, Zeang Sheng, Meiyi Qiang, Yang Li, Wentao Zhang

机构 * University of Electronic Science and Technology of China(电子科技大学) Peking University(北京大学) Tencent Inc(腾讯公司) Zhongguancun Academy(中关村学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);image-text(abstract);分类 cs.CV

AI总结 研究揭示多模态大语言模型在多模态检索中表现不佳的原因,通过稀疏自编码器分析发现其表示空间主要由文本语义构成,视觉语义不足,导致检索性能下降,提出ReAlign方法提升检索效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08884 2023-10-16 cs.CV 87%

Extending Multi-modal Contrastive Representations

Zehan Wang, Ziang Zhang, Luping Liu, Yang Zhao, Haifeng Huang, Tao Jin, Zhou Zhao

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);audio-visual(abstract)

Comments Our code is available at https://github.com/MCR-PEFT/Ex-MCR

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01426 2022-07-05 cs.MM cs.AI cs.CL cs.CV 87%

Dynamic Contrastive Distillation for Image-Text Retrieval

Jun Rao, Liang Ding, Shuhan Qi, Meng Fang, Yang Liu, Li Shen, Dacheng Tao

专题命中 跨模态检索 :image-text(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05174 2023-10-12 cs.IR cs.CV cs.LG cs.MM 87%

Scene-centric vs. Object-centric Image-Text Cross-modal Retrieval: A Reproducibility Study

Mariya Hendriksen, Svitlana Vakulenko, Ernst Kuiper, Maarten de Rijke

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(title);分类 cs.CV、cs.MM

Comments 18 pages, accepted as a reproducibility paper at ECIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14910 2021-10-01 cs.CV cs.AI cs.LG 87%

CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations

Mohammadreza Zolfaghari, Yi Zhu, Peter Gehler, Thomas Brox

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(title);分类 cs.CV、cs.AI

Comments ICCV 2021, 14 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08108 2021-04-19 cs.CV cs.CL 87%

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

Shir Gur, Natalia Neverova, Chris Stauffer, Ser-Nam Lim, Douwe Kiela, Austin Reiter

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00670 2020-05-05 cs.LG cs.CL cs.CV cs.HC stat.ML 87%

Stochastic Neighbor Embedding of Multimodal Relational Data for Image-Text Simultaneous Visualization

Morihiro Mizutani, Akifumi Okuno, Geewook Kim, Hidetoshi Shimodaira

专题命中 跨模态检索 :multimodal(title,abstract);image-text(title);分类 cs.CV、cs.CL

Comments 20 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15906 2026-06-16 cs.IR cs.AI cs.CL cs.DB cs.MM 新提交 87%

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA

MAGE-RAG:面向长文档问答的多粒度自适应图证据多模态RAG

Yilong Zuo, Xunkai Li, Jing Yuan, Qiangqiang Dai, Hongchao Qin, Ronghua Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.MM

AI总结 提出MAGE-RAG框架,通过离线构建包含页面和元素节点的证据图,在线自适应构建证据子图,平衡证据覆盖与噪声控制,在长文档多模态问答中取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02571 2025-02-25 cs.CL cs.AI cs.CV cs.IR cs.LG 87%

MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, Wei Ping

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at ICLR 2025. We release the model weights at: https://huggingface.co/nvidia/MM-Embed

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24114 2024-11-01 cs.CV cs.AI cs.CL 87%

Nearest Neighbor Normalization Improves Multimodal Retrieval

Neil Chowdhury, Franklin Wang, Sumedh Shenoy, Douwe Kiela, Sarah Schwettmann, Tristan Thrush

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04681 2024-07-08 cs.CV cs.AI cs.CL cs.LG 87%

Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

Yuanze Lin, Yunsheng Li, Dongdong Chen, Weijian Xu, Ronald Clark, Philip Torr, Lu Yuan

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16150 2025-11-21 cs.CV 86%

Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval

基于推理的嵌入:利用多模态大语言模型推理提升多模态检索

Chunxu Liu, Jiyuan Yang, Ruopeng Gao, Yuhan Zhu, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title);分类 cs.CV

AI总结 本文提出基于推理的嵌入方法,利用多模态大语言模型的推理能力提升多模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11001 2025-02-18 q-bio.BM cs.AI cs.LG q-bio.QM 86%

CL-MFAP: A Contrastive Learning-Based Multimodal Foundation Model for Molecular Property Prediction and Antibiotic Screening

Gen Zhou, Sugitha Janarthanan, Yutong Lu, Pingzhao Hu

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

Comments Gen Zhou and Sugitha Janarthanan contributed equally; Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16166 2023-05-26 cs.CL 86%

Multimodal Relation Extraction with Cross-Modal Retrieval and Synthesis

Xuming Hu, Zhijiang Guo, Zhiyang Teng, Irwin King, Philip S. Yu

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title);分类 cs.CL

Comments Accepted to ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06489 2021-12-14 cs.CV 86%

Multi-Modal Mutual Information Maximization: A Novel Approach for Unsupervised Deep Cross-Modal Hashing

Tuan Hoang, Thanh-Toan Do, Tam V. Nguyen, Ngai-Man Cheung

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16829 2024-09-02 astro-ph.HE astro-ph.IM cs.LG 86%

Maven: A Multimodal Foundation Model for Supernova Science

Gemma Zhang, Thomas Helfer, Alexander T. Gagliano, Siddharth Mishra-Sharma, V. Ashley Villar

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title)

Comments code: https://github.com/ThomasHelfer/multimodal-supernovae data: https://huggingface.co/datasets/thelfer/multimodal_supernovae

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02148 2026-08-04 cs.IR cs.CL cs.CV 新提交 86%

Douyin Multimodal Embedding Model Technical Report

抖音多模态嵌入模型技术报告

Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL

AI总结 针对现有多模态嵌入模型难以兼顾效率与细粒度区分的问题,提出分两阶段训练的DME模型,在MMEB-v2数据集及抖音生产场景中均取得优异效果。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17832 2026-05-28 cs.LG cs.AI cs.CR cs.CV 86%

MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks

MM-PoisonRAG:通过局部和全局投毒攻击破坏多模态RAG

Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios, Saikrishna Sanniboina, Nanyun Peng, Kai-Wei Chang, Daniel Kang, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出MM-PoisonRAG框架,通过局部投毒攻击(LPA)和全局投毒攻击(GPA)两种策略,系统研究多模态检索增强生成(RAG)在知识投毒下的脆弱性,实验表明攻击成功率高达56%且能绕过现有防御。

Comments Code is available at https://github.com/HyeonjeongHa/MM-PoisonRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10120 2026-05-12 cs.CV cs.AI 86%

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

MicroWorld: 通过多模态属性图赋能多模态大语言模型,弥合微观领域差距

Manyu Li, Ruian He, Chenxi Ma, Weimin Tan, Bo Yan

机构 * Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理关键实验室) School of Computer Science, Fudan University(复旦大学计算机科学学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 MicroWorld通过构建多模态属性图增强多模态大语言模型在微观领域的能力,无需领域微调即可提升推理性能,实验显示在MicroVQA和MicroBench基准上均取得显著提升。

Comments 29 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27934 2026-05-01 cs.AI cs.CL 86%

MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

MM-StanceDet:基于检索的多模态多智能体立场检测

Weihai Lu, Zhejun Zhao, Yanshu Li, Huan He

机构 * Peking University(北京大学) Baidu Inc(百度公司) Brown University(布朗大学) Amazon(亚马逊)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MM-StanceDet框架,通过整合检索增强、多模态分析代理、辩论阶段和自我反思,解决多模态立场检测中的上下文 grounding、跨模态解释模糊和单次推理脆弱问题,实验表明其在五个数据集上优于现有方法。

Comments Accepted on ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05318 2026-04-28 cs.CV cs.AI 86%

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA

mKG-RAG:利用多模态知识图谱在检索增强生成中进行知识密集型视觉问答

Xu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan, Qing Li

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 本文提出mKG-RAG框架,通过整合多模态知识图谱提升检索增强生成在知识密集型视觉问答中的性能,采用双阶段检索策略和图提取方法构建高质量知识图谱,实验表明优于现有方法。

Comments In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'26), July 20-24, 2026, Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17054 2026-04-21 cs.CV cs.AI 86%

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

mEOL:无需训练的指令引导多模态嵌入器用于矢量图形和图像检索

Kyeong Seon Kim, Baek Seong-Eun, Lee Jung-Mok, Tae-Hyun Oh

机构 * KAIST(韩国科学技术院) POSTECH

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 本文提出无需训练的指令引导多模态嵌入框架,通过多模态大语言模型将文本、位图和SVG代码映射到对齐的嵌入空间,利用模态特定指令和结构化SVG提示实现嵌入方向控制,构建首个文本到SVG检索基准,展示出优于传统基线的性能。

Comments Round 1 early acceptance to WACV 2026, Project page: https://scene-the-ella.github.io/meol

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05826 2025-12-04 cs.CV cs.AI cs.LG 86%

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

从像素到 prose:推进遥感多模态语言模型

Xintian Sun, Benji Peng, Charles Zhang, Fei Jin, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Xinyuan Song, Yichao Zhang

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Minnesota - Twin Cities(明尼苏达大学双城分校) Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文探讨了多模态语言模型在遥感中的应用,分析了其技术基础、挑战及未来发展方向,强调了其在环境监测和灾害响应中的重要作用。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18072 2025-09-04 eess.IV cs.AI cs.CV 86%

Multimodal Medical Image Binding via Shared Text Embeddings

Yunhao Liu, Suyang Xi, Shiqi Liu, Hong Ding, Chicheng Jin, Chong Zhong, Junjun He, Catherine C. Liu, Yiqing Shen

机构 * The Hong Kong Polytechnic University(香港理工大学) Emory University(埃默里大学) The University of Hong Kong(香港大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) Johns Hopkins University(约翰霍普金斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14824 2025-06-19 cs.LG cs.AI cs.MM 86%

FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models

Yao Zhang, Hewei Gao, Haokun Chen, Weiguo Li, Yunpu Ma, Volker Tresp

机构 * LMU Munich(慕尼黑莱布尼茨大学) Technical University of Munich(慕尼黑技术大学) Siemens Technology(西门子技术) Munich Center for Machine Learning(慕尼黑机器学习中心) University Heidelberg(海德堡大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06269 2025-04-10 cs.IR cs.AI cs.CL 86%

EXCLAIM: An Explainable Cross-Modal Agentic System for Misinformation Detection with Hierarchical Retrieval

Yin Wu, Zhengxuan Zhang, Fuling Wang, Yuyu Luo, Hui Xiong, Nan Tang

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00330 2025-01-03 cs.CL cs.AI cs.IR 86%

Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion

Hebin Wang, Yangning Li, Yinghui Li, Hai-Tao Zheng, Wenhao Jiang, Hong-Gee Kim

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13480 2024-10-27 cs.CV cs.IR cs.MM 86%

A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels

Haochen Han, Minnan Luo, Huan Liu, Fang Nan

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06644 2024-09-12 cs.CV cs.AI 86%

EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis

Danli Shi, Weiyi Zhang, Jiancheng Yang, Siyu Huang, Xiaolan Chen, Mayinuer Yusufu, Kai Jin, Shan Lin, Shunming Liu, Qing Zhang, Mingguang He

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏