arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3433 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3433 篇

2604.15628 2026-04-20 cs.CV cs.CL cs.IR cs.LG cs.MM 93%

SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding

SIMMER:通过基于MLLM的嵌入进行跨模态食物图像-食谱检索

Keisuke Gomi, Keiji Yanai

机构 * The University of Electro-Communications(电通大学)

专题命中 跨模态检索 :MLLM(title,title_cn);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本文提出SIMMER,利用基于MLLM的嵌入模型VLM2Vec,通过统一编码器处理食物图像和食谱文本,设计定制化提示模板和数据增强策略,实现跨模态检索的高精度。

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14766 2026-07-30 cs.CV cs.CL 版本更新 92%

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

ASCD:用于减少多模态大语言模型(MLLM)幻觉的注意力可导向对比解码

Yujun Wang, Aniri, Jinhe Bi, Soeren Pirk, Yunpu Ma

专题命中 跨模态检索 :MLLM(title,title_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 该研究针对MLLM的幻觉问题,提出ASCD方法,通过正负引导调整解码时的注意力分数,在多基准上显著减少幻觉并提升VQA准确率,且无需额外训练。

Comments Accepted at AAAI 2026

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306-10314, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03174 2025-06-05 cs.CV cs.AI cs.LG 91%

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

Koki Matsuishi, Kosuke Ukita, Tsuyoshi Okita

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(title);分类 cs.CV、cs.AI

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02791 2026-08-05 cs.CV 新提交 91%

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

更好、更强、更快、更广泛:基于多模态大型语言模型(MLLM)的结构化全掩码预测分割

Jiazhen Liu, Mingkuan Feng, Long Chen

机构 * The Hong Kong University of Science and Technology (HKUST)(香港科技大学)

专题命中 跨模态检索 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV

AI总结 STAMPlus通过结构化全掩码预测解耦自回归对话与非自回归掩码预测,解决了MLLM分割的三难问题,在提升性能的同时降低延迟,实现多类开放词汇等分割任务的SOTA表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05909 2026-05-08 cs.AI 91%

Null Space Constrained Contrastive Visual Forgetting for MLLM Unlearning

空域约束对比遗忘用于MLLM反学习

Yuhang Wang, Zhenxing Niu, Haoxuan Ji, Guangyu He, Linlin Zhang, Haichang Gao

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Xi’an Jiaotong University(西安交通大学)

专题命中 跨模态检索 :MLLM(title,title_cn);multimodal(abstract);分类 cs.AI

AI总结 本文提出一种MLLM反学习方法,通过冻结LLM主干并微调视觉模块,在保留非目标视觉知识和全部文本知识的同时,有效遗忘目标视觉知识。

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25273 2026-04-29 cs.CV 90%

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

对抗大多模态模型中的视觉忽视与语义漂移以提升跨模态检索

Guosheng Zhang, Linkai Liu, Keyao Wang, Haixiao Yue, Zhiwen Tan, Xiao Tan

机构 * Baidu Inc(百度公司)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出SSA-ME框架,通过显式建模显著视觉主体,提升细粒度表示学习,解决多模态检索中的语义漂移和视觉模态忽视问题,实验表明其在MMEB基准上达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10805 2024-02-19 cs.MM cs.AI cs.CL cs.CV cs.IR 90%

Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond

Yongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie, Wenjie Li, Tat-Seng Chua

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10016 2020-02-28 cs.IR cs.AI cs.CL cs.CV cs.LG 90%

Deep Multimodal Image-Text Embeddings for Automatic Cross-Media Retrieval

Hadi Abdi Khojasteh, Ebrahim Ansari, Parvin Razzaghi, Akbar Karimi

专题命中 跨模态检索 :multimodal(title,abstract);image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 6 pages and 2 figures, Learn more about this project at https://iasbs.ac.ir/~ansari/deeptwitter

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19663 2025-12-23 cs.CV cs.AI 90%

Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis

超越CLIP:基于知识的多模态Transformer用于糖尿病视网膜病变诊断中的跨模态对齐

Argha Kamal Samanta, Harshika Goyal, Vasudha Joshi, Tushar Mungle, Pabitra Mitra

机构 * Department of Medicine Stanford University Stanford, USA(医学系 斯坦福大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种基于知识的多模态Transformer框架,通过整合视网膜图像、临床文本和结构化数据,提升糖尿病视网膜病变诊断中的跨模态对齐与检索性能。

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19961 2024-10-01 cs.CV cs.CL 90%

Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval

Yabing Wang, Le Wang, Qiang Zhou, Zhibin Wang, Hao Li, Gang Hua, Wei Tang

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(title);multi-modal(abstract);MLLM(abstract)

Comments Accepted by ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15670 2023-09-06 cs.CV cs.AI 90%

Multimodal Foundation Models For Echocardiogram Interpretation

Matthew Christensen, Milos Vukadinovic, Neal Yuan, David Ouyang

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14057 2023-08-14 cs.LG cs.AI cs.CL 90%

Cross-modal Contrastive Learning for Multimodal Fake News Detection

Longzheng Wang, Chuang Zhang, Hongbo Xu, Yongxiu Xu, Xiaohan Xu, Siqi Wang

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10824 2023-04-24 cs.CV cs.MM 90%

Rethinking Benchmarks for Cross-modal Image-text Retrieval

Weijing Chen, Linli Yao, Qin Jin

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted to SIGIR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15694 2026-06-16 cs.MM cs.AI cs.CV cs.LG 新提交 89%

MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs

MAF: 面向情感分析的多模态自适应少样本提示方法

Hangling Xie

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 跨模态检索 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 提出MAF框架,通过动态检索与查询相关的多模态示例,利用轻量级系数生成网络实时融合多模态相似度,结合多数投票提升MLLM在情感分析中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02329 2026-03-04 cs.CV 89%

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

HAMMER: 通过跨模态整合利用大语言模型进行意图驱动的3D affordance grounding

Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau

机构 * The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :MLLM(title,abstract);cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 HAMMER通过跨模态整合多模态大语言模型,实现意图驱动的3D affordance grounding,提升3D表示的准确性和鲁棒性。

Comments Accepted by CVPR 2026. Project Page: https://rayyoh.github.io/Hammer

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20624 2026-02-25 cs.AI cond-mat.stat-mech 89%

Physics-based phenomenological characterization of cross-modal bias in multimodal models

基于物理现象的多模态模型跨模态偏差表征

Hyeongmo Kim, Sohyun Kang, Yerin Choi, Seungyeon Ji, Junhyuk Woo, Hyunsuk Chung, Soyeon Caren Han, Kyungreem Han

机构 * B rain Science Institute(脑科学研究院) Korea Institute of Science and Technology(韩国科学技术院) Department of Physics and Astronomy(物理与天文学系) Department of Computer Science and Engineering(计算机科学与工程系) University of Science and Technology KIST School(科学技术KIST学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出基于物理现象的多模态模型跨模态偏差表征方法,揭示多模态输入可能强化模态主导性。

Comments Best Paper Award at BiasinAI track in AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13684 2024-05-24 cs.CL 89%

CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models

Guangzhi Sun, Potsawee Manakul, Adian Liusie, Kunat Pipatanakul, Chao Zhang, Phil Woodland, Mark Gales

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);audio-visual(abstract);分类 cs.CL

Comments 21 pages. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10029 2024-05-20 cs.MM 89%

AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(title,abstract);multimodal(abstract);分类 cs.MM

Comments This work has been strong-accepted as the oral conference paper by IEEE International Conference on Multimedia & Expo (ICME) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08789 2023-07-19 cs.CV 89%

Efficient Token-Guided Image-Text Retrieval with Consistent Multimodal Contrastive Training

Chong Liu, Yuqi Zhang, Hongsong Wang, Weihua Chen, Fan Wang, Yan Huang, Yi-Dong Shen, Liang Wang

专题命中 跨模态检索 :multimodal(title,abstract);image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Code is publicly available: https://github.com/LCFractal/TGDT

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05814 2022-10-13 cs.LG cs.CV 89%

SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval

Minyoung Kim

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05481 2022-09-14 cs.LG cs.AI 89%

A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language

Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, Ji-Rong Wen

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04712 2026-05-12 cs.CV cs.AI eess.IV 89%

SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation

SAR-RAG:通过语义搜索、检索和MLLM生成实现目标识别的视觉问答

David F. Ramirez, Tim Overman, Kristen Jaskie, Joe Marvin, Andreas Spanias

机构 * SenSIP Center, School of ECEE, Arizona State University(SenSIP中心,电子与计算机工程学院,亚利桑那州立大学) Prime Solutions Group Inc(Prime Solutions Group公司)

专题命中 跨模态检索 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SAR-RAG方法,结合多模态大语言模型和语义嵌入向量数据库,通过语义搜索和检索提升SAR图像目标识别的准确性,通过分类和回归指标验证效果。

Comments Accepted to 2026 SPIE Defense + Security, Automatic Target Recognition XXXVI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07841 2023-03-28 cs.CV cs.AI cs.MM 89%

Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

Guangxing Han, Long Chen, Jiawei Ma, Shiyuan Huang, Rama Chellappa, Shih-Fu Chang

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02601 2021-12-07 cs.IR 89%

Variational Autoencoder with CCA for Audio-Visual Cross-Modal Retrieval

Jiwei Zhang, Yi Yu, Suhua Tang, Jianming Wu, Wei Li

专题命中 跨模态检索 :cross-modal(title,abstract);audio-visual(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16161 2026-06-16 cs.CV 新提交 89%

Multimodal LLM-Empowered Re-Ranking for Generalizable Person Re-Identification

多模态大语言模型赋能的通用行人重识别重排序

Jiachen Li, Xiaojin Gong

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院)

专题命中 跨模态检索 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 提出利用多模态大语言模型(MLLM)的泛化能力,通过微调MLLM并计算μ-距离来改进推理阶段的重排序,从而提升领域泛化行人重识别的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15979 2026-04-20 cs.CV 89%

MMGait: Towards Multi-Modal Gait Recognition

MMGait:迈向多模态步态识别

Chenye Wang, Qingyuan Cai, Saihui Hou, Aoqi Li, Yongzhen Huang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

专题命中 跨模态检索 :multi-modal(title,summary_cn);cross-modal(abstract);分类 cs.CV

AI总结 本文提出MMGait多模态步态基准,整合五种异构传感器数据,评估单模态、跨模态和多模态步态识别方法,引入Omni Multi-Modal Gait Recognition任务和OmniGait基线模型。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15134 2026-06-16 cs.CV cs.AI cs.LG 新提交 88%

Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

超越标量距离:来自冻结MLLM的语义属性梯度用于视觉嵌入

Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :MLLM(title_cn,summary_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出SAGA框架,利用冻结的多模态大语言模型(MLLM)通过GRPO奖励机制为视觉编码器提供属性级监督,替代传统标量距离,提升零样本图像检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12824 2025-12-16 cs.CV cs.AI 88%

Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners

为少样本学习适应多模态基础模型:对比captioners的全面研究

N. K. B. M. P. K. B. Narasinghe, Uthayasanker Thayasivam

机构 * Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka(计算机科学与工程系,穆塔瓦大学,斯里兰卡)

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文研究了如何通过对比captioners适应少样本学习,探讨了参数高效微调策略及生成-对比基础模型的高效适应方法。

Comments 9 pages, 3 figures. Accepted to VISAPP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13515 2025-12-09 cs.CV cs.AI 88%

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

UniME-V2:基于大规模语言模型的通用多模态嵌入学习判别器

Tiancheng Gu, Kaicheng Yang, Kaichen Zhang, Xiang An, Ziyong Feng, Yueyi Zhang, Weidong Cai, Jiankang Deng, Lidong Bing

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV、cs.AI

AI总结 UniME-V2利用大规模语言模型提升通用多模态嵌入学习的判别能力,通过引入MLLM-as-a-Judge机制和重排序模型实现更高效的负样本挖掘和语义匹配。

Comments AAAI2026 Oral, Webpage:https://garygutc.github.io/UniME-v2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23895 2025-09-30 cs.CV cs.AI 88%

Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios

Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu Li

机构 * Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏