arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2307.12996 2023-07-26 cs.LG cs.AI cs.CL cs.IR q-bio.QM 76%

Extracting Molecular Properties from Natural Language with Multimodal Contrastive Learning

Romain Lacombe, Andrew Gaut, Jeff He, David Lüdeke, Kateryna Pistunova

专题命中 跨模态检索 :multimodal(title);分类 cs.CL、cs.AI

Comments 2023 ICML Workshop on Computational Biology

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10893 2023-04-24 cs.CV cs.MM 76%

FindVehicle and VehicleFinder: A NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system

Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Xiaohui Zhu, Jeremy Smith, Eng Gee Lim, Yutao Yue

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12413 2022-12-27 cs.CV cs.AI 76%

Segmentation of Parotid Gland Tumors Using Multimodal MRI and Contrastive Learning

Zi'an Xu, Yin Dai, Fayu Liu, Boyuan Wu, Weibing Chen, Lifu Shi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14395 2022-10-27 cs.CV cs.CL cs.LG 76%

IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Alireza Dirafzoon, Aparajita Saraf, Amy Bearman, Babak Damavandi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01445 2022-03-07 cs.CV cs.CL cs.LG eess.IV 76%

LILE: Look In-Depth before Looking Elsewhere -- A Dual Attention Network using Transformers for Cross-Modal Information Retrieval in Histopathology Archives

Danial Maleki, H. R Tizhoosh

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12932 2019-10-01 cs.CV cs.HC cs.IR cs.MM 76%

BUDA.ART: A Multimodal Content-Based Analysis and Retrieval System for Buddha Statues

Benjamin Renoust, Matheus Oliveira Franca, Jacob Chan, Van Le, Ayaka Uesaka, Yuta Nakashima, Hajime Nagahara, Jueren Wang, Yutaka Fujioka

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments Demo video at: https://www.youtube.com/watch?v=3XJvLjSWieY

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.11299 2019-05-15 cs.CV cs.CL 76%

Image search using multilingual texts: a cross-modal learning approach between image and text

Maxime Portaz, Hicham Randrianarivo, Adrien Nivaggioli, Estelle Maudet, Christophe Servan, Sylvain Peyronnet

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.00860 2017-07-27 cs.LG cs.AI cs.CV 76%

Conditional generation of multi-modal data using constrained embedding space mapping

Subhajit Chaudhury, Sakyasingha Dasgupta, Asim Munawar, Md. A. Salam Khan, Ryuki Tachibana

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.AI

Comments 7 pages, 4 figures, ICML 2017 Workshop on Implicit Models

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.07831 2017-05-18 cs.RO cs.AI cs.CV cs.LG 76%

Deep Multimodal Embedding: Manipulating Novel Objects with Point-clouds, Language and Trajectories

Jaeyong Sung, Ian Lenz, Ashutosh Saxena

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments IEEE International Conference on Robotics and Automation (ICRA), 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.06078 2016-04-15 cs.CV cs.CL cs.LG 76%

Learning Deep Structure-Preserving Image-Text Embeddings

Liwei Wang, Yin Li, Svetlana Lazebnik

专题命中 跨模态检索 :image-text(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04812 2026-03-03 cs.CV cs.AI cs.CL cs.LG 75%

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

LLaVE: 大规模语言和视觉嵌入模型与基于难度加权的对比学习

Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 LLaVE通过基于难度加权的对比学习提升多模态嵌入模型性能,实现SOTA表现和强泛化能力。

Comments Accepted by Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17960 2025-10-22 astro-ph.IM astro-ph.CO 75%

AION-1: Omnimodal Foundation Model for Astronomical Sciences

Liam Parker, Francois Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Nathan Casserau, Pierre Cornette, Keiya Hirashima, Geraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Bruno Regaldo-Saint Blancard, Kyunghyun Cho, Miles Cranmer, Shirley Ho

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)

Comments Accepted at Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23379 2025-10-20 cs.CL cs.AI cs.CV 75%

CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(计算科学学院,格拉斯哥大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint, 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12131 2025-03-18 cs.CV cs.AI cs.LG cs.SD eess.AS 75%

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu

专题命中 跨模态检索 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15052 2025-01-28 cs.CV cs.AI cs.MM 75%

Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval

Bingjun Luo, Jinpeng Wang, Wang Zewen, Junjie Zhu, Xibin Zhao

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00986 2022-12-06 cs.CV cs.AI cs.CL 75%

Masked Contrastive Pre-Training for Efficient Video-Text Retrieval

Fangxun Shu, Biaolong Chen, Yue Liao, Shuwen Xiao, Wenyu Sun, Xiaobo Li, Yousong Zhu, Jinqiao Wang, Si Liu

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07194 2020-08-05 cs.CL cs.CV cs.IR cs.LG cs.MM 75%

Recommending Themes for Ad Creative Design via Visual-Linguistic Representations

Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, Wei Wang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments 7 pages, 8 figures, 2 tables, accepted by The Web Conference 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14841 2026-08-18 cs.AI 新提交 74%

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

重排序器所见:面向长文档多模态问答的多方面页面标注

Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee, Carl Yang

机构 * Emory University(埃默里大学) Hippocratic AI(希波克拉底人工智能公司)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

AI总结 本文针对长文档多模态问答的重排序瓶颈,提出含Trident-R与Trident-S组件的Trident模型,通过多方面页面标注提升检索与生成性能,在多数据集上取得显著效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07291 2026-08-10 cs.CV 新提交 74%

Foundation Models Adaptation for Multi-View Multi-modal Cardiac MRI Segmentation and Direct Ejection Fraction Estimation

基础模型在多视图多模态心脏MRI分割与直接射血分数估算中的适配

Sina Amirrajab, Cian M Scannell, Volker Vehof, Michael Bietenbeck, Ali Yilmaz

机构 * Maastricht University(马斯特里赫特大学) University Hospital Münster(明斯特大学医院) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

AI总结 本研究针对CMR-Multi挑战,微调CineMA并组合冻结CMR基础模型,实现多视图CMR分割与直接LVEF估算,验证了基础模型适配多视图CMR分析的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00544 2026-08-04 cs.CV 新提交 74%

Zero-Cost Virtual RNA: Approximating Immunotherapy Signatures via Cross-Modal WSI Retrieval

零成本虚拟RNA:通过跨模态WSI检索近似免疫治疗特征

Sigrid Vila-Bagaria, Mar Teixidó, Miquel Piñol, Felip Vilardell, Robert Montal, Veronica Vilaplana

机构 * Universitat Politècnica de Catalunya(加泰罗尼亚理工大学) IRB Lleida(莱里达生物医学研究所)

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

AI总结 针对胃腺癌免疫治疗RNA特征检测成本高的问题,提出VITA模型,通过H&E与RNA跨模态对齐实现仅用H&E切片近似RNA特征,取得0.72准确率等结果,提供低成本预筛选工具。

Comments Accepted to MIDL 2026 Short Paper track

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15152 2026-07-30 cs.CV 74%

Caption-Matching: A Multimodal Approach for Cross-Domain Image Retrieval

图像跨域检索:一种多模态方法

Lucas Iijima, Nikos Giakoumoglou, Tania Stathaki

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院电气与电子工程系)

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

AI总结 本文提出Caption-Matching方法,利用预训练视觉语言模型生成的图像描述作为中间表示,实现跨域图像检索,无需标注数据或进一步训练,在Office-Home和DomainNet上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10698 2026-07-14 cs.LG cs.CV 新提交 74%

On the modality gap and the contrastive loss in multi-modal representation learning

关于多模态表示学习中的模态差距和对比损失

Fabian Mager, Hiba Nassar, Lars Kai Hansen

机构 * Technical University of Denmark(丹麦技术大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

AI总结 研究CLIP风格双编码器对比学习中的模态差距,发现是InfoNCE公式在低温下的模式失败所致。提出xNCE方法,使用跨模态和模态内负对比对,能缩小模态差距,提升零样本分类性能且不牺牲判别几何。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04339 2026-07-07 cs.LG cs.AI cs.CR 新提交 74%

One Framework for All: Cross-Modal Membership Inference for Generative Models

一框架适用于所有:生成模型的跨模态成员推理

Dayong Ye, Tainqing Zhu, Kun Gao, Junhao Liu, Yichuan Chen, Shuai Zhou, Hengzhu Liu, Bo Liu, Wanlei Zhou

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

AI总结 研究生成模型的跨模态成员推理问题,基于生成模型输出分布可近似训练数据分布的特性,在共享嵌入空间建模,通过似然比测试推理,贡献统一框架,性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02605 2026-06-03 cs.LG cs.AI eess.IV 74%

Cross-Modal Contrastive Learning of ECG and Angiography Representations for Severe Stenosis Classification

用于严重狭窄分类的心电图与血管造影表示的跨模态对比学习

Nikola Cenikj, Özgün Turgut, Alexander Müller, Alexander Steger, Jan Kehrer, Marcus Brugger, Daniel Rueckert, Philip Müller

机构 * Chair for AI in Healthcare and Medicine, Technical University of Munich and TUM University Hospital(人工智能在医疗与医学中的研究所,慕尼黑技术大学及慕尼黑大学医院) Department of Computing, Imperial College London(伦敦帝国理工学院计算机系) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML)) Department of Internal Medicine, TUM University Hospital(慕尼黑大学医院内科学系)

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

AI总结 提出StenCE预训练框架,通过跨模态对比学习从心电图特征中实现冠状动脉狭窄风险分层,在严重狭窄分类中首次达到高性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20460 2026-04-23 cs.CV 74%

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

CCTVBench:用于多模态大语言模型的对比一致性交通视频问答基准

Xingcheng Zhou, Hao Guo, Rui Song, Walter Zimmer, Mingyu Liu, André Schamschurko, Hu Cao, Alois Knoll

机构 * Technical University of Munich(慕尼黑技术大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

AI总结 CCTVBench通过真实事故视频与世界模型生成的反事实场景配对,提出对比一致性评估方法,揭示视频问答中对比一致性与标准指标间的显著差距,并引入C-TCD方法提升问答与一致性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05474 2026-03-16 cs.IR cs.AI 74%

LLM-driven Multimodal Recommendation

基于大语言模型的多模态推荐

Yicheng Di

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

AI总结 本文提出LMMRec框架,通过建模用户动机提升推荐系统的可解释性和说服力,实验表明其在三个真实数据集上效果显著。

Comments There are some writing errors in our methods section that need to be corrected. We will then add extensive experiments and rewrite the Introduction and related work sections

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06698 2026-03-11 cs.CV 74%

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

突破几何瓶颈:不对称跨模态蒸馏中的对比扩展

Kabir Thayani

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

AI总结 本研究通过对比扩展方法解决不对称跨模态蒸馏中的维度坍缩问题,揭示了容量与密度之间的权衡关系。

Comments Introduced auxiliary InfoNCE objective to reverse dimensional collapse. Expanded experiments to DINOv2 teacher and CIFAR-100 dataset. 3 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12883 2026-02-16 eess.IV cs.CV 74%

Dual-Phase Cross-Modal Contrastive Learning for CMR-Guided ECG Representations for Cardiovascular Disease Assessment

双相跨模态对比学习用于CMR引导的ECG表示以评估心血管疾病

Laura Alvarez-Florez, Angel Bujalance-Gomez, Femke Raijmakers, Samuel Ruiperez-Campillo, Maarten Z. H. Kolk, Jesse Wiers, Julia Vogt, Erik J. Bekkers, Ivana Išgum, Fleur V. Y. Tjong

机构 * Amsterdam University Medical Center(阿姆斯特丹大学医学中心) University of Amsterdam(阿姆斯特丹大学) ETH Zurich(苏黎世联邦理工学院) Department of Radiology and Nuclear Medicine, Amsterdam University Medical Center(放射医学与核医学系,阿姆斯特丹大学医学中心) Department of Radiology, Mayo Clinic, Rochester(放射医学系,梅奥诊所,罗切斯特)

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

AI总结 本文提出双相跨模态对比学习方法,通过结合ECG和CMR数据提升心血管疾病评估的ECG表示能力。

Comments Paper accepted at SPIE Medical Imaging 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11101 2025-10-13 cs.CV cs.LG 74%

A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis

Asifullah Khan, Laiba Asmatullah, Anza Malik, Shahzaib Khan, Hamna Asif

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 38 pages, 8 figures, survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11557 2025-10-09 cs.CR cs.AI 74%

AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs

Lara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas Cavigelli

专题命中 跨模态检索 :multi-modal(title);分类 cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏