arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2009.01485 2021-10-22 cs.CV cs.AI 73%

SAC: Semantic Attention Composition for Text-Conditioned Image Retrieval

Surgan Jandial, Pinkesh Badjatiya, Pranit Chawla, Ayush Chopra, Mausoom Sarkar, Balaji Krishnamurthy

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Surgan Jandial, Pinkesh Badjatiya, Pranit Chawla, and Ayush Chopra contributed equally to this work. Work accepted at WACV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08872 2021-10-19 cs.CV cs.IR cs.LG cs.MM 73%

Contrastive Learning of Visual-Semantic Embeddings

Anurag Jain, Yashaswi Verma

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08688 2021-08-20 cs.CL cs.CV 73%

Contrastive Language-Image Pre-training for the Italian Language

Federico Bianchi, Giuseppe Attanasio, Raphael Pisoni, Silvia Terragni, Gabriele Sarti, Sri Lakshmi

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06247 2021-05-14 cs.CL cs.CV cs.IR 73%

Video Corpus Moment Retrieval with Contrastive Learning

Hao Zhang, Aixin Sun, Wei Jing, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments 11 pages, 7 figures and 6 tables. Accepted by SIGIR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10054 2021-04-21 cs.CV cs.MM 73%

T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval

Xiaohan Wang, Linchao Zhu, Yi Yang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted to CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02949 2020-10-08 cs.CV cs.CL cs.LG 73%

Learning to Represent Image and Text with Denotation Graph

Bowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie, Fei Sha

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments to appear at EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.02971 2019-11-11 cs.CL cs.CV cs.LG 73%

Probing Contextualized Sentence Representations with Visual Awareness

Zhuosheng Zhang, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Hai Zhao

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06514 2019-10-16 cs.CV cs.MM 73%

Target-Oriented Deformation of Visual-Semantic Embedding Space

Takashi Matsubara

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.05612 2018-07-31 cs.LG cs.CL cs.CV 73%

VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, Sanja Fidler

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted as spotlight presentation at British Machine Vision Conference (BMVC) 2018. Code: https://github.com/fartashf/vsepp

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.01720 2018-04-09 cs.CV cs.CL cs.LG 73%

Finding beans in burgers: Deep semantic-visual embedding with localization

Martin Engilberge, Louis Chevallier, Patrick Pérez, Matthieu Cord

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to CVPR2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02123 2026-03-09 cs.CL 72%

CMRAG: Co-modality-based visual document retrieval and question answering

CMRAG:基于共模的视觉文档检索与问答

Wang Chen, Wenhan Yu, Guanqiang Qi, Weikang Li, Yang Li, Lei Sha, Deguo Xia, Jizhou Huang

机构 * Baidu Inc(百度公司) The University of Hong Kong(香港大学) Beihang University(北京航空航天大学) Peking University(北京大学)

专题命中 跨模态检索 :multimodal(abstract,comments);cross-modal(abstract);分类 cs.CL

AI总结 CMRAG通过统一编码模型和共模信息检索方法,提升多模态视觉文档问答系统的性能。

Comments Published at ICLR 2026 Workshop on Multimodal Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17203 2026-08-19 stat.ML cs.LG math.ST stat.TH 新提交 71%

Expressivity In Multimodal Contrastive Learning

多模态对比学习中的表达能力

Andrew Stuart, Florian Wolf

专题命中 跨模态检索 :multimodal(title)

AI总结 该研究针对多模态对比学习的表达能力展开分析,明确CLIP架构的表达能力随模态数量变化,提出Hadamard-CLIP模型,实现任意数量模态联合分布的通用近似且保留CLIP的检索优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19733 2026-07-14 physics.chem-ph cs.LG 版本更新 71%

NMIRacle: Multi-modal Generative Molecular Elucidation from IR and NMR Spectra

NMIRacle:从红外和核磁共振谱进行多模生成分子解析

Federico Ottomano, Yingzhen Li, Alex M. Ganose

机构 * Federico Ottomano Yingzhen Li Alex M. Ganose

专题命中 跨模态检索 :multi-modal(title)

AI总结 NMIRacle通过两阶段生成框架,利用红外和核磁共振光谱数据实现分子结构的高精度解析,优于现有方法并保持复杂性下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15514 2026-06-16 cs.RO cs.LG 新提交 71%

Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities

强化学习引导的软融合检索用于缺失模态下的鲁棒多模态模仿学习

Hassan Ismkhan, Hamid Bouchahcia

机构 * Bournemouth University(伯恩茅斯大学)

专题命中 跨模态检索 :multimodal(title)

AI总结 提出RL4IL方法,利用强化学习策略从训练库中检索最相关专家演示,并通过软交叉注意力融合生成动作,有效处理传感器缺失问题,在LIBERO基准上超越现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11382 2026-06-11 cs.LG q-bio.BM 新提交 71%

GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction

GLACIER:用于分子性质预测的多模态师生基础模型

Emily Nguyen, Yongchan Hong, Harsh Toshniwal, Yan Liu, Andreas Luttens

机构 * Department of Computer Science, University of Southern California(南加州大学计算机科学系) Department of Quantitative and Computational Biology, University of Southern California(南加州大学定量与计算生物学系) Amazon(亚马逊) Department of Medical Biochemistry and Biophysics, Science for Life Laboratory, Karolinska Institutet(卡罗林斯卡学院医学生物化学与生物物理系,生命科学实验室)

专题命中 跨模态检索 :multimodal(title)

AI总结 提出GLACIER师生框架,通过融合分子图、SMILES和物理化学描述符三种模态,并利用大模型蒸馏,实现高效准确的分子性质预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08997 2026-05-12 eess.SP 71%

SEM-RAG: Structure-Preserving Multimodal Graph Compilation and Entropy-Guided Retrieval for Telecommunication Standards

SEM-RAG:结构保持的多模态图编译与熵引导检索用于电信标准

Yuzhi Yang, Lina Bariah, Yuhuan Lu, Hang Zou, Mérouane Debbah

专题命中 跨模态检索 :multimodal(title)

AI总结 本文提出SEM-RAG框架,通过结构保持编译和熵引导检索提升电信标准文档检索性能,实验显示其在表格和公式密集型问题上表现优异,准确率高达94.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11028 2026-01-19 cs.LG 71%

AVP-Pro: An Adaptive Multi-Modal Fusion and Contrastive Learning Approach for Comprehensive Two-Stage Antiviral Peptide Identification

AVP-Pro: 一种自适应多模态融合与对比学习方法用于全面两阶段抗病毒肽鉴定

Xinru Wen, Weizhong Lin, zi liu, Xuan Xiao

专题命中 跨模态检索 :multi-modal(title)

AI总结 AVP-Pro通过自适应多模态融合和对比学习方法,实现了抗病毒肽的高效两阶段鉴定,提升了模型判别能力和小样本下的分类精度。

Comments arXiv admin note: substantial text overlap with arXiv:2512.21544

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21544 2025-12-29 cs.LG 71%

AVP-Fusion: Adaptive Multi-Modal Fusion and Contrastive Learning for Two-Stage Antiviral Peptide Identification

AVP-Fusion:适应性多模态融合与对比学习用于两阶段抗病毒肽识别

Xinru Wen, Weizhong Lin, Xuan Xiao

专题命中 跨模态检索 :multi-modal(title)

AI总结 AVP-Fusion通过自适应多模态融合与对比学习,实现了抗病毒肽的两阶段高效识别,显著提升准确率和决策边界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01558 2025-10-03 cs.CE cs.LG eess.SP 71%

CardioRAG: A Retrieval-Augmented Generation Framework for Multimodal Chagas Disease Detection

Zhengyang Shen, Xuehao Zhai, Hua Tu, Mayue Shi

机构 * Department of Electrical and Electronic Engineering Imperial College London(帝国理工学院电子与电气工程系) Department of Civil and Environmental Engineering Imperial College London(帝国理工学院土木与环境工程系) Institute of Biomedical Engineering Department of Engineering Science University of Oxford(牛津大学生物医学工程研究所)

专题命中 跨模态检索 :multimodal(title)

Comments 4 pages, 2 figures. Accepted for oral presentation at the 52nd international Computing in Cardiology Conference (CinC2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16768 2025-06-23 cs.IR 71%

eSapiens: A Real-World NLP Framework for Multimodal Document Understanding and Enterprise Knowledge Processing

Isaac Shi, Zeyuan Li, Wenli Wang, Lewei He, Yang Yang, Tianyu Shi

专题命中 跨模态检索 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08953 2025-05-15 q-bio.QM 71%

Multimodal Modeling of Ultradian Rhythms Using the Hankel Alternative View of Koopman (HAVOK) Analysis

Emmanuel Molefi, Billy C. Smith, Christopher Thornton, Peter N. Taylor, Yujiang Wang

专题命中 跨模态检索 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11251 2024-12-03 cs.IR 71%

Unifying Multimodal Retrieval via Document Screenshot Embedding

Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, Jimmy Lin

专题命中 跨模态检索 :multimodal(title)

Comments EMNLP2024 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20325 2024-10-29 cs.LG cs.SI 71%

Domain Specific Data Distillation and Multi-modal Embedding Generation

Sharadind Peddiraju, Srini Rajagopal

专题命中 跨模态检索 :multi-modal(title)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00955 2024-08-06 cs.LG cs.CY 71%

FairEHR-CLP: Towards Fairness-Aware Clinical Predictions with Contrastive Learning in Multimodal Electronic Health Records

Yuqing Wang, Malvika Pillai, Yun Zhao, Catherine Curtin, Tina Hernandez-Boussard

专题命中 跨模态检索 :multimodal(title)

Comments MLHC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15151 2024-04-24 q-bio.NC 71%

Semantic distance organizes social knowledge: Insights from semantic dementia and cross-modal conceptual space

Y. Ivette Colón, Matthew Rouse, Matthew A. Lambon Ralph, Timothy T. Rogers

专题命中 跨模态检索 :cross-modal(title)

Comments 6 pages, CogSci proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06556 2023-08-15 cs.IR 71%

Contrastive Learning for Cross-modal Artist Retrieval

Andres Ferraro, Jaehun Kim, Sergio Oramas, Andreas Ehmann, Fabien Gouyon

专题命中 跨模态检索 :cross-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11080 2023-04-24 eess.SP cs.LG 71%

Multimodal contrastive learning for diagnosing cardiovascular diseases from electrocardiography (ECG) signals and patient metadata

Tue M. Cao, Nhat H. Tran, Phi Le Nguyen, Hieu Pham

专题命中 跨模态检索 :multimodal(title)

Comments Accepted for presentation at the Midwest Machine Learning Symposium (MMLS 2023), Chicago, IL, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06991 2023-04-17 cs.IR 71%

WYTIWYR: A User Intent-Aware Framework with Multi-modal Inputs for Visualization Retrieval

Shishi Xiao, Yihan Hou, Cheng Jin, Wei Zeng

专题命中 跨模态检索 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14779 2022-11-29 cs.CR cs.LG q-fin.ST 71%

Who is Gambling? Finding Cryptocurrency Gamblers Using Multi-modal Retrieval Methods

Zhengjie Huang, Zhenguang Liu, Jianhai Chen, Qinming He, Shuang Wu, Lei Zhu, Meng Wang

专题命中 跨模态检索 :multi-modal(title)

Journal ref International Journal of Multimedia Information Retrieval (2022): 1-13

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12866 2019-12-02 cs.IR 71%

Macross: Urban Dynamics Modeling based on Metapath Guided Cross-Modal Embedding

Yunan Zhang, Heting Gao, Tarek Abdelzaher

专题命中 跨模态检索 :cross-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏