arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2603.11520 2026-03-13 cs.CV cs.AI 84%

FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval

FBCIR: 在组合图像检索中平衡跨模态聚焦

Chenchen Zhao, Jianhuan Zhuo, Muxi Chen, Zhaohua Zhang, Wenyu Jiang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Qiang Xu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出FBCIR方法,通过分析CIR模型中的聚焦不平衡问题,提出数据增强工作流程以提升模型在挑战性场景下的性能。

Comments 20 pages, 5 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06265 2026-03-09 cs.CV cs.AI 84%

SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability

SPARC:概念对齐的稀疏自编码器用于跨模型和跨模态可解释性

Ali Nasiri-Sarvi, Hassan Rivaz, Mahdi S. Hosseini

机构 * Department of Computer Science and Software Engineering (CSSE) Concordia University, Canada(计算机科学与软件工程系(CSSE)康科迪亚大学,加拿大) Department of Electrical and Computer Engineering (ECE) Concordia University, Canada(电气与计算机工程系(ECE)康科迪亚大学,加拿大)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 SPARC通过稀疏自编码器实现跨模型和跨模态的概念对齐,提升概念表示的一致性与可解释性

Comments Accepted at TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13704 2026-03-06 cs.IR cs.AI cs.CV 84%

Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search

Pailitao-VL:统一的嵌入与重排序用于实时多模态工业搜索

Lei Chen, Chen Ju, Xu Chen, Zhicheng Wang, Yuheng Jiao, Hongfeng Zhan, Zhaoyang Li, Shihao Xu, Zhixiang Zhao, Tong Jia, Lin Li, Yuan Gao, Jun Song, Jinsong Lan, Xiaoyong Zhu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 Pailitao-VL通过统一嵌入与重排序方法,解决多模态工业搜索中的高精度、实时性和抗噪声问题,实现先进检索架构的高效部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23589 2026-03-03 cs.CV cs.AI 84%

Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models

伪对比学习用于多模态模型中的图表理解

Hiroshi Sasaki

机构 * The Japan Research Institute, Limited(日本研究机构有限公司)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出伪对比学习方法,通过生成合成图表增强多模态模型对图表的理解能力,实验表明在图像-文本匹配和视觉问答任务中优于传统CLIP方法。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00510 2026-03-03 cs.CV cs.AI 84%

What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models

视觉标记究竟编码了什么?揭示多模态大语言模型中的稀疏性与冗余性

Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究揭示多模态大语言模型中视觉标记的稀疏性与冗余性,通过EmbedLens工具发现活跃标记在进入模型前已编码细粒度信息,并提出中层注入方法提升效率。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01959 2026-03-02 cs.CV cs.AI cs.LG 84%

Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models

结构感知对比学习用于多模态模型的图表理解

Hiroshi Sasaki

机构 * The Japan Research Institute, Limited(日本研究机构)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出结构感知对比学习方法,通过专门的损失函数提升多模态模型对图表图像的理解能力,在图像-文本匹配和视觉问答任务中取得显著改进。

Comments 10 pages, 8 figures

Journal ref ICCVW (2025) 7463-7472

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07909 2026-02-16 cs.LG cs.AI cs.CV 84%

Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning

解释并缓解对比多模态学习中的模态差距

Can Yaras, Siyi Chen, Peng Wang, Qing Qu

机构 * Department of Electrical Engineering \& Computer Science, University of Michigan

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文通过分析对比多模态学习中的模态差距成因,提出通过温度调度和模态交换等策略缓解模态差距,从而提升多模态任务性能。

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04988 2026-02-13 cs.CV cs.AI 84%

Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model

遥感检索增强生成:通过多模态数据集和检索增强生成模型连接遥感图像与综合知识

Congcong Wen, Yiting Lin, Xiaokang Qu, Nan Li, Yong Liao, Xiang Li, Hui Lin

机构 * School of Cyber Science and Technology, University of Science and Technology of China(信息科学技术学院,中国科学技术大学) China Academy of Electronics and Information Technology(电子信息技术研究院)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出RS-RAG框架,通过多模态数据集和检索增强生成模型,提升遥感图像与综合知识的连接能力,有效提升复杂查询的语义推理性能。

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine (GRSM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04195 2026-02-10 cs.CV cs.CL 84%

Cross-Modal Retrieval for Motion and Text via DropTriple Loss

通过DropTriple损失实现动与文本的跨模态检索

Sheng Yan, Yang Liu, Haoqiang Wang, Xin Du, Mengyuan Liu, Hong Liu

机构 * School of Artificial Intelligence, Chongqing University of Technology, China(重庆理工大学人工智能学院) College of Computer Science, Sichuan University, China(四川大学计算机学院) Key Laboratory of Machine Perception, Shenzhen Graduate School, Peking University, China(北京大学深圳研究生院机器感知重点实验室)

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 本文提出DropTriple损失,用于提升人类动作与文本之间的跨模态检索性能,实验表明在HumanML3D数据集上实现了较高的检索准确率。

Comments This paper has been accepted by ACM MM Asia 2023 (Best Paper Candidate)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08099 2026-02-10 cs.CV cs.AI 84%

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec:解锁视频MLLM嵌入用于视频-文本检索

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 VidVec通过利用预训练MLLM的中间层嵌入和校准头部,实现无需训练的视频-文本检索,超越现有方法,达到最佳性能。

Comments Project page: https://iyttor.github.io/VidVec/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07125 2026-02-10 cs.IR cs.AI cs.CV cs.LG 84%

Reasoning-Augmented Representations for Multimodal Retrieval

增强推理的表示用于多模态检索

Jianrui Zhang, Anirudh Sundara Rajan, Brandon Han, Soochahn Lee, Sukanta Ganguly, Yong Jae Lee

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Kookmin University(韩国高丽大学)

专题命中 跨模态检索 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出一种增强推理的多模态检索方法,通过外部化推理和语义密集表示提升检索性能,尤其在知识密集型查询和组合修改请求中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05648 2026-01-12 q-bio.GN cs.AI cs.CL cs.LG 84%

Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training

开放世界知识辅助的单细胞基础模型与鲁棒跨模态细胞-语言预训练

Haoran Wang, Xuanyi Zhang, Shuangsang Fang, Longke Ran, Ziqing Deng, Yong Zhang, Yuxiang Li, Shaoshuai Li

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 OKR-CELL通过开放世界知识和鲁棒跨模态预训练,提升单细胞基础模型在细胞-语言跨模态任务中的表现。

Comments 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05399 2026-01-12 cs.CV cs.AI cs.IR 84%

Multi-task Cross-modal Learning for Chest X-ray Image Retrieval

多任务跨模态学习用于胸部X射线图像检索

Zhaohui Liang, Sivaramakrishnan Rajaraman, Niccolo Marini, Zhiyun Xue, Sameer Antani

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出多任务学习框架,通过改进BiomedCLIP模型,提升胸部X射线图像与文本的跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24064 2026-01-01 cs.CV cs.MM 84%

Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal Retrieval

具有噪声标签的邻居感知实例细化用于跨模态检索

Yizhi Liu, Ruitao Pu, Shilin Xu, Yingke Chen, Quan-Hui Liu, Yuan Sun

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

AI总结 本文提出了一种新的跨模态学习框架NIRNL,通过引入噪声标签和邻居感知机制,提升模型在高噪声环境下的检索性能。

Comments 9 pages, 4 figures, and AAAI-26 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14885 2025-12-10 cs.CV cs.CL 84%

You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction

你可能可以自由发言:通过答案提取提升多模态大语言模型的细粒度视觉识别能力

Logan Lawrence, Oindrila Saha, Megan Wei, Chen Sun, Subhransu Maji, Grant Van Horn

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校) Brown University(布朗大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 本文提出nlg2choice方法,通过两阶段策略提升多模态大语言模型在细粒度视觉识别任务中的表现,通过开放性问题和受限解码提高检索效率。

Comments Accepted to WACV26. 12 pages, 8 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21827 2025-12-01 cs.AI cs.CV 84%

Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI

评估合成临床笔记的策略以用于医疗多模态AI

Niccolo Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue, Sameer Antani

机构 * Division of Intramural Research, National Library of Medicine, National Institutes of Health(国家医学图书馆内部研究部,国家卫生研究院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了合成临床笔记的生成策略,通过优化提示设计和医学元数据,提升医疗多模态AI在分类和跨模态检索任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15435 2025-11-20 cs.CV cs.AI cs.IR 84%

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

HV-Attack:多模态检索增强生成的分层视觉攻击

Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, Hanjiang Lai

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-Sen University(中山大学) Singapore Management University(新加坡管理学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种分层视觉攻击方法,通过在图像输入中添加不可察觉扰动,破坏多模态检索增强生成系统的检索和生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13545 2025-11-18 cs.CV cs.AI 84%

Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks

Md. Iqbal Hossain, Afia Sajeeda, Neeresh Kumar Perla, Ming Shao

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10892 2025-11-17 cs.CV cs.AI 84%

MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition

Feng Li, Ke Wu, Yongwei Li

机构 * School of Computer and Information Engineering, Anhui University of Finance and Economics, Anhui, 233030, China(计算机与信息工程学院,安徽财经大学,安徽,233030,中国) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences, Beijing, 100045, China(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院,北京,100045,中国)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by 32nd International Conference on MultiMedia Modeling (MMM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16470 2025-11-10 cs.IR cs.CL cs.CV 84%

Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering

Kuicai Dong, Yujing Chang, Shijie Huang, Yasheng Wang, Ruiming Tang, Yong Liu

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Paper accepted to NeurIPS 2025 DB

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23617 2025-10-29 cs.LG cs.AI cs.CL 84%

An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis

Phuong Q. Dao, Mark Roantree, Vuong M. Ngo

机构 * Ho Chi Minh City Open University, Ho Chi Minh, Vietnam Insight Centre for Data Analytics \& School of Computing, Dublin City University, Dublin, Ireland

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments The paper has been accepted for presentation at the MEDES 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01622 2025-10-03 cs.IR cs.AI cs.CL 84%

LLM4Rec: Large Language Models for Multimodal Generative Recommendation with Causal Debiasing

Bo Ma, Hang Li, ZeHua Hu, XiaoFan Gui, LuYao Liu, Simon Lau

机构 * Department of Software \& Microelectronics Peking University Beijing, China Department of Software \& Microelectronics Peking University Beijing, China hangli\ Department of Software \& Microelectronics Peking University Beijing, China zehua\ Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ Economic Law School China University of Political Science Peking University Changsha, China

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12653 2025-09-17 cs.CV cs.AI 84%

Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations

Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong

机构 * Hefei University of Technology(合肥工业大学) University of Trento(特伦托大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09721 2025-09-15 cs.CV cs.AI cs.LG 84%

A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval

Jiayi Miao, Dingxin Lu, Zhuqi Wang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07879 2025-09-15 cs.IR cs.AI cs.CV 84%

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

Wei Yang, Jingjing Fu, Rui Wang, Jinyu Wang, Lei Song, Jiang Bian

机构 * Microsoft Research Asia(微软亚洲研究院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22896 2025-08-01 cs.HC cs.AI cs.CV cs.RO 84%

iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement

Kohou Wang, ZhaoXiang Liu, Lin Bai, Kun Fan, Xiang Liu, Huan Hu, Kai Wang, Shiguo Lian

机构 * Unicom Data Intelligence(中国联通数据智能研究所) Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院) China United Network Communications Group Corporation Limited(中国联合网络通信集团有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18902 2025-07-08 cs.AI cs.CL cs.IR 84%

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval

Michael Günther, Saba Sturua, Mohammad Kalim Akram, Isabelle Mohr, Andrei Ungureanu, Bo Wang, Sedigheh Eslami, Scott Martens, Maximilian Werk, Nan Wang, Han Xiao

机构 * Jina AI GmbH(Jina AI公司)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 1-10 main, 14-22 experimental results, benchmark tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04590 2025-07-08 cs.CV cs.CL 84%

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Rui Meng, Ziyan Jiang, Ye Liu, Mingyi Su, Xinyi Yang, Yuepeng Fu, Can Qin, Zeyuan Chen, Ran Xu, Caiming Xiong, Yingbo Zhou, Wenhu Chen, Semih Yavuz

机构 * Salesforce Research(Salesforce研究机构) UC Santa Barbara(加州大学圣巴巴拉分校) University of Waterloo(滑铁卢大学) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19650 2025-05-28 cs.CV cs.IR cs.MM 84%

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Fanheng Kong, Jingyuan Zhang, Yahui Liu, Hongzhi Zhang, Shi Feng, Xiaocui Yang, Daling Wang, Yu Tian, Victoria W., Fuzheng Zhang, Guorui Zhou

机构 * Northeastern University(东北大学) Kuaishou Technology(快手科技)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 26 pages, project page: https://friedrichor.github.io/projects/UNITE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13023 2025-04-18 cs.CL cs.CV 84%

ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images

Sangwook Kim, Soonyoung Lee, Jongseong Jang

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏