arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2506.21538 2025-06-27 cs.CV cs.IR cs.LG 83%

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Hani Alomari, Anushka Sivakumar, Andrew Zhang, Chris Thomas

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17782 2025-06-24 cs.IR cs.AI 83%

Expanding Relevance Judgments for Medical Case-based Retrieval Task with Multimodal LLMs

Catarina Pires, Sérgio Nunes, Luís Filipe Teixeira

机构 * Faculty of Engineering, University of Porto(葡萄牙波尔图大学工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments To appear at the Third Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2025), co-located with SIGIR 2025. 9 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05929 2025-06-19 cs.LG cs.AI 83%

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

Hongyang Lei, Xiaolong Cheng, Qi Qin, Dan Wang, Kun Fan, Huazhen Huang, Qingqing Gu, Yetao Wu, Zhonglin Jiang, Yong Chen, Luo Ji

机构 * Geely AI Lab(Geely人工智能实验室) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Peking University(北京大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 5 figures. ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11674 2025-06-16 cs.CV 83%

Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports

Libin Lan, Hongxing Li, Zunhui Xia, Juan Zhou, Xiaofei Zhu, Yongmei Li, Yudong Zhang, Xin Luo

机构 * College of Computer Science and Engineering, Chongqing University of Technology(重庆理工大学计算机科学与工程学院)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE TMI for possible publication. Our code is available at https://github.com/violet-42/CM-CGNS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11184 2025-06-04 cs.CL cs.AI cs.CV cs.MM 83%

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu

机构 * Renmin University of China(中国人民大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Chinese University of Hong Kong(香港中文大学) Johns Hopkins University(约翰霍普金斯大学) Hong Kong University of Science and Technology(香港科技大学) Tencent(腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02020 2025-06-04 cs.CV cs.LG 83%

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Youze Xue, Dian Li, Gang Liu

机构 * Tencent(腾讯)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24361 2025-06-02 cs.CV 83%

Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation

Roger Ferrod, Cássio F. Dantas, Luigi Di Caro, Dino Ienco

机构 * University of Turin(都灵大学) INRAE, UMR TETIS, Univ. Montpellier(法国蒙彼利埃大学、INRAE、UMR TETIS) EVERGREEN, Univ. Montpellier, Inria(法国蒙彼利埃大学、EVERGREEN、Inria)

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19707 2025-05-27 cs.CV cs.IR 83%

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval

Rong-Cheng Tu, Zhao Jin, Jingyi Liao, Xiao Luo, Yingjie Wang, Li Shen, Dacheng Tao

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院) Department of Computer Science, University of California, Los Angeles, USA(加州大学洛杉矶分校计算机科学系) Sun Yat-sen University Shenzhen Campus, School of Cyber Science and Technology, Shenzhen, China(中山大学深圳校区信息科学与技术学院)

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02087 2025-05-06 cs.AI 83%

Retrieval-augmented in-context learning for multimodal large language models in disease classification

Zaifu Zhan, Shuang Zhou, Xiaoshan Zhou, Yongkang Xiao, Jun Wang, Jiawen Deng, He Zhu, Yu Hou, Rui Zhang

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 17 Pages, 1 figure, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15344 2025-05-06 cs.SD eess.AS 83%

Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions

Yifei Xin, Yuexian Zou

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 eess.AS

Comments Accepted by Interspeech2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21044 2025-05-01 cs.CR cs.AI 83%

AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection

Jianbo Gao, Keke Gai, Jing Yu, Liehuang Zhu, Qi Wu

机构 * Beijing Institute of Technology(北京理工大学) School of Information Engineering, Minzu University of China(民族大学信息工程学院) University of Adelaide(阿德莱德大学)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02698 2025-04-29 cs.LG cs.AI q-bio.QM 83%

SCMPPI: Supervised Contrastive Multimodal Framework for Predicting Protein-Protein Interactions

Shengrui XU, Tianchi Lu, Zikun Wang, Jixiu Zhai

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 20 pages,9 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16576 2025-04-24 cs.IR cs.AI 83%

MMHCL: Multi-Modal Hypergraph Contrastive Learning for Recommendation

Xu Guo, Tong Zhang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学) Shituoyun (Nanjing) Technology Co., Ltd(石墨云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 23 pages, 8 figures. This manuscript is currently under major revision for ACM Transactions on Multimedia Computing, Communications, and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10074 2025-04-22 cs.AI 83%

MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework

Zihan Ling, Zhiyao Guo, Yixuan Huang, Yi An, Shuai Xiao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02264 2025-04-04 cs.CV 83%

MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception

Wenzhuo Liu, Wenshuo Wang, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Pengfei Li, Zilong Chen, Huiming Yang, Zhiwei Li, Lening Wang, Tiao Tan, Huaping Liu

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01476 2025-04-03 cs.CV 83%

Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction

Junlong Ren, Hao Wang

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14977 2025-04-03 cs.AI eess.IV 83%

Trustworthy Enhanced Multi-view Multi-modal Alzheimer's Disease Prediction with Brain-wide Imaging Transcriptomics Data

Shan Cong, Zhoujie Fan, Hongwei Liu, Yinghan Zhang, Xin Wang, Haoran Luo, Xiaohui Yao

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16855 2025-04-02 cs.CL cs.IR 83%

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to CVPR 2025, models at https://huggingface.co/Alibaba-NLP/gme-Qwen2-VL-2B-Instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12287 2025-03-24 cs.CL 83%

CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model

Dongyoung Go, Taesun Whang, Chanhee Lee, Hwa-Yeon Kim, Sunghoon Park, Seunghwan Ji, Jinho Kim, Dongchan Kim, Young-Bum Kim

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12914 2025-03-18 cs.CV 83%

Efficient Multimodal 3D Object Detector via Instance-Level Contrastive Distillation

Zhuoqun Su, Huimin Lu, Shuaifeng Jiao, Junhao Xiao, Yaonan Wang, Xieyuanli Chen

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06277 2025-03-18 cs.CV 83%

STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classification

Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 16 pages (including 5 pages of supplementary materials), accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18363 2025-03-12 cs.CV 83%

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Qing Jiang, Gen Luo, Yuqin Yang, Yuda Xiong, Yihao Chen, Zhaoyang Zeng, Tianhe Ren, Lei Zhang

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 35 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15268 2025-02-07 cs.CL 83%

Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation

Liwen Sun, James Zhao, Megan Han, Chenyan Xiong

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CL

Comments NAACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07157 2025-01-14 cs.AI 83%

CureGraph: Contrastive Multi-Modal Graph Representation Learning for Urban Living Circle Health Profiling and Prediction

Jinlin Li, Xiao Zhou

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01720 2024-12-03 cs.CV 83%

LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

Yikun Liu, Pingan Chen, Jiayin Cai, Xiaolong Jiang, Yao Hu, Jiangchao Yao, Yanfeng Wang, Weidi Xie

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00373 2024-12-03 cs.LG cs.AI math.AG 83%

Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment

Dongfang Zhao

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00341 2024-12-03 cs.CV eess.IV 83%

Fusing Physics-Driven Strategies and Cross-Modal Adversarial Learning: Toward Multi-Domain Applications

Hana Satou, Alan Mitkiy

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05756 2024-11-06 cs.CV 83%

GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification

Souhail Bakkali, Sanket Biswas, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol, Oriol Ramos Terrades, Josep Lladós

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16977 2024-10-23 cs.CL 83%

IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing

Kang Chen, Qingheng Zhang, Chengbao Lian, Yixin Ji, Xuwei Liu, Shuguang Han, Guoqiang Wu, Fei Huang, Jufeng Chen

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10115 2024-10-16 cs.CV cs.LG cs.RO 83%

Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection

Mehar Khurana, Neehar Peri, James Hays, Deva Ramanan

专题命中 跨模态检索 :cross-modal(title);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments The first two authors contributed equally. This work has been accepted to the Conference on Robot Learning (CoRL) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏