arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2101.08533 2022-06-14 cs.CV 74%

Eliminate Deviation with Deviation for Data Augmentation and a General Multi-modal Data Learning Method

Yunpeng Gong, Liqing Huang, Lifei Chen

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04426 2022-05-26 cs.CV 74%

Deep Feature Rotation for Multimodal Image Style Transfer

Son Truong Nguyen, Nguyen Quang Tuyen, Nguyen Hong Phuc

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Accepted to NICS'21

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15914 2022-05-24 cs.CV 74%

Tasting the cake: evaluating self-supervised generalization on out-of-distribution multimodal MRI data

Alex Fedorov, Eloy Geenjaar, Lei Wu, Thomas P. DeRamus, Vince D. Calhoun, Sergey M. Plis

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Presented as a RobustML workshop paper at ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04021 2022-05-02 cs.CV 74%

Supervised Contrastive Learning for Detecting Anomalous Driving Behaviours from Multimodal Videos

Shehroz S. Khan, Ziting Shen, Haoying Sun, Ax Patel, Ali Abedi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 8 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07052 2022-04-15 cs.CV 74%

CroCo: Cross-Modal Contrastive learning for localization of Earth Observation data

Wei-Hsin Tseng, Hoàng-Ân Lê, Alexandre Boulch, Sébastien Lefèvre, Dirk Tiede

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments Accepted for publication in the ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences (online from July 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05431 2021-11-11 cs.LG cs.AI 74%

Multi-Task Prediction of Clinical Outcomes in the Intensive Care Unit using Flexible Multimodal Transformers

Benjamin Shickel, Patrick J. Tighe, Azra Bihorac, Parisa Rashidi

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11592 2021-10-25 cs.CV cs.IR 74%

Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering

Zhongwei Xie, Ling Liu, Yanzhao Wu, Luo Zhong, Lin Li

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments accepted by ACM Transactions on Information Systems(TOIS). arXiv admin note: text overlap with arXiv:2108.00705, arXiv:2108.03788

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.01654 2021-08-12 cs.CV 74%

Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query

Guanyu Cai, Jun Zhang, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Pai Peng, Xiaowei Guo, Feiyue Huang, Xing Sun

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments Accepted by ICCV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13711 2021-06-28 cs.IR cs.CL 74%

Multimodal Emergent Fake News Detection via Meta Neural Process Networks

Yaqing Wang, Fenglong Ma, Haoyu Wang, Kishlay Jha, Jing Gao

专题命中 跨模态检索 :multimodal(title);分类 cs.CL

Comments accepted by KDD 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08665 2021-05-19 cs.LG cs.CV 74%

A multimodal deep learning framework for scalable content based visual media retrieval

Ambareesh Ravi, Amith Nandakumar

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Paper pertaining to a course project

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12482 2021-02-23 cs.CV cs.SI 74%

On the Limits to Multi-Modal Popularity Prediction on Instagram -- A New Robust, Efficient and Explainable Baseline

Christoffer Riis, Damian Konrad Kowalczyk, Lars Kai Hansen

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments Presented at ICAART 2021

Journal ref Proceedings of the 13th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART, ISBN 978-989-758-484-8, pages 1200-1209, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.04152 2020-01-07 cs.LG cs.IR cs.MM stat.ML 74%

Learning Discriminative Hashing Codes for Cross-Modal Retrieval based on Multi-view Features

Jun Yu, Xiao-Jun Wu, Josef Kittler

专题命中 跨模态检索 :cross-modal(title);分类 cs.MM

Comments 28 pages, 10 figures, 13 tables. The paper is under consideration at Pattern Analysis and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02013 2019-08-07 cs.CV 74%

Generalised Zero-Shot Learning with a Classifier Ensemble over Multi-Modal Embedding Spaces

Rafael Felix, Ben Harwood, Michele Sasdelli, Gustavo Carneiro

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.08389 2018-07-31 cs.CV 74%

Conditional Image-Text Embedding Networks

Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng, Robinson Piramuthu, Svetlana Lazebnik

专题命中 跨模态检索 :image-text(title);分类 cs.CV

Comments ECCV 2018 accepted paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03134 2018-05-09 cs.CV 74%

Image Retrieval with Mixed Initiative and Multimodal Feedback

Nils Murrugarra-Llerena, Adriana Kovashka

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments In submission to BMVC 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.03470 2018-05-03 cs.CV 74%

Learning Two-Branch Neural Networks for Image-Text Matching Tasks

Liwei Wang, Yin Li, Jing Huang, Svetlana Lazebnik

专题命中 跨模态检索 :image-text(title);分类 cs.CV

Comments accepted version in TPAMI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1207.1522 2012-07-09 cs.CV cs.NE 74%

Multimodal similarity-preserving hashing

Jonathan Masci, Michael M. Bronstein, Alexander A. Bronstein, Jürgen Schmidhuber

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12843 2026-08-14 cs.CV cs.AI 新提交 73%

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

用于基于文本的行人异常检索的、带分歧感知重排序的异构视觉语言集成方法

Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen

机构 * Hanoi University of Science and Technology(河内科技大学) University of Information Technology, VNU-HCM(胡志明市国家大学信息技术大学) Vietnam National University, Ho Chi Minh City(胡志明市国家大学) VinUniversity Vietnamese-German University(越南德国大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 该研究针对大规模基于文本的行人异常检索难题,提出带分歧感知重排序的异构视觉语言集成方法,在PAB基准上取得90.92% mAP等优异指标,验证了方案有效性。

Comments Accepted at the ECCV 2026 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22919 2026-07-28 cs.CV cs.AI 新提交 73%

Controlling Embedding Spaces with Text-Conditioned Transformations

通过文本条件变换控制嵌入空间

Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah, Kushal Kafle

机构 * Institute of Artificial Intelligence, University of Central Florida(中佛罗里达大学人工智能研究所) Adobe Research(Adobe研究院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 研究提出文本条件变换视觉嵌入,通过自然语言描述属性类别,网络生成仿射变换强调指定属性,能同时学习多属性,可在推理时访问,为控制嵌入空间提供统一高效框架,在多任务中展现近零推理成本的最优性能。

Comments Accepted at ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03728 2026-05-11 cs.CV cs.AI 73%

CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval

CSMCIR: 基于记忆库的增强对称对齐用于复合图像检索

Zhipeng Qian, Zihan Liang, Yufei Ma, Ben Chen, Huangyu Dai, Yiwei Ma, Jiayi Ji, Chenyi Lei, Han Li, Xiaoshuai Sun

机构 * Kuaishou Technology(快播科技) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CSMCIR,通过多级链式思维提示策略、对称双塔架构和熵基记忆库策略,实现高效查询-目标对齐,提升复合图像检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08508 2026-04-23 cs.CV cs.CL 73%

Re:Verse -- Can Your VLM Read a Manga?

Re:Verse -- 能读懂漫画吗?

Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学) Indian Institute of Technology, Jodhpur(印度理工学院,朱达浦尔) Indian Institute of Technology, Varanasi(印度理工学院,瓦拉纳西)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文通过分析漫画叙事理解,揭示现有VLM在时间因果和跨面板连贯性上的不足,提出新的评估框架,系统研究长篇叙事理解能力。

Comments Accepted (oral) at ICCV (AISTORY Workshop) 2025

Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 3820-3830

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14204 2026-04-17 cs.SD cs.AI eess.AS 73%

Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition

解耦双分支图学习用于对话情感识别

Chengling Guo, Yuntao Shou, Tao Meng, Wei Ai, Yun Tan, Keqin Li

机构 * College of Computer and Mathematics(计算机与数学学院) Central South University of Forestry and Technology(林业科技大学) Changsha Hospital for Maternal and Child Health Care(长沙妇幼保健医院) Department of Computer Science, State University of New York(纽约州立大学计算机科学系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出解耦双分支图学习框架,通过分离模态不变和特定特征,结合傅里叶图神经网络和说话人感知超图,提升对话情感识别性能。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13777 2026-04-14 cs.CV cs.AI cs.SD 73%

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping

Sat2Sound:一种用于零样本声音景观映射的统一框架

Subash Khanal, Srikumar Sastry, Aayush Dhakal, Adeel Ahmad, Abby Stylianou, Nathan Jacobs

机构 * Washington University in St. Louis(圣路易斯华盛顿大学) Taylor Geospatial(泰勒地理空间) Saint Louis University(圣路易斯大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 Sat2Sound通过融合多模态数据,实现地球表面声音分布的零样本映射,通过对比学习和代码本对齐学习发现跨模态的'声音景观概念',在GeoSound和SoundingEarth基准上取得最佳性能。

Comments Accepted to EarthVision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06638 2026-03-24 cs.CV cs.AI 73%

StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering

StaR-KVQA:用于隐式知识视觉问答的结构化推理轨迹

Zhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu, Hui Zhang, Weicheng Zhu, Xin Zhang

机构 * Ant International, Ant Group(蚂蚁集团国际部,蚂蚁集团) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院) Anhui Provincial Key Laboratory of High Performance Computing(安徽省高性能计算重点实验室)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 StaR-KVQA通过引入双路径结构化推理轨迹提升隐式知识视觉问答的准确性与推理透明度,采用自蒸馏方法构建轨迹增强数据集,无需外部检索工具,在OK-VQA基准上实现11.3%的精度提升。

Comments 8+3+3 pages, code: https://github.com/jianyingzhihe/StaR-KVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08973 2026-03-20 cs.CV cs.AI 73%

Is Contrastive Distillation Enough for Learning Comprehensive 3D Representations?

对比蒸馏是否足以学习全面的3D表示?

Yifan Zhang, Junhui Hou

机构 * School of Mechatronic Engineering and Automation, Shanghai University, Shanghai, China(上海大学机械与自动化工程学院) Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China(香港城市大学计算机科学系)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出CMCR框架,通过整合模态共享和特定特征,改进传统方法,引入掩码图像建模和占用估计任务,提升3D表示学习效果。

Comments Accepted to IJCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10889 2026-03-05 cs.CV cs.AI cs.LG 73%

Topological Alignment of Shared Vision-Language Embedding Space

共享视觉-语言嵌入空间的拓扑对齐

Junwon You, Dasol Kang, Jae-Hun Jung

机构 * Department of Mathematics, POSTECH(POSTECH数学系) BootCamp, Google(Google BootCamp)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 ToMCLIP通过拓扑对齐方法提升多语言视觉-语言模型的结构一致性和零样本性能

Comments 27 pages, 5 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02434 2026-03-04 cs.CV cs.AI 73%

MIRAGE: Knowledge Graph-Guided Cross-Cohort MRI Synthesis for Alzheimer's Disease Prediction

MIRAGE:基于知识图谱的跨群体MRI合成用于阿尔茨海默病预测

Guanchen Wu, Zhe Huang, Yuzhang Xie, Runze Yan, Akul Chopra, Deqiang Qiu, Xiao Hu, Fei Wang, Carl Yang

机构 * Emory University, Atlanta, USA(埃默里大学) Cornell University, New York, USA(康奈尔大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MIRAGE通过知识图谱引导的跨群体MRI合成,提升无真实MRI群体中阿尔茨海默病的诊断准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20095 2026-03-03 cs.CV cs.CL cs.LG 73%

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

BioCAP: 利用合成描述性注释超越标签在生物基础模型中的应用

Ziheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Samuel Stevens, Hilmar Lapp, Tanya Berger-Wolf, Yu Su, Wei-Lun Chao, Jianyang Gu

机构 * The Ohio State University(俄亥俄州立大学) Duke University(杜克大学) Boston University(波士顿大学)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

AI总结 BioCAP通过生成合成描述性注释提升生物基础模型的性能,实现物种分类和文本-图像检索的高准确率。

Comments ICLR 2026; Project page: https://imageomics.github.io/biocap/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17687 2026-02-23 cs.IR cs.AI cs.CL cs.LG 73%

IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering

IRPAPERS: 一个用于科学检索和问答的视觉文档基准

Connor Shorten, Augustas Skaburskas, Daniel M. Jones, Charles Pierse, Roberto Esposito, John Trengrove, Etienne Dilocker, Bob van Luijt

机构 * Weaviate

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

AI总结 IRPAPERS基准通过比较图像和文本检索系统,揭示了多模态混合搜索在科学文档检索和问答中的优势。

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13961 2026-02-17 cs.CV astro-ph.IM cs.CL 73%

MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

MarsRetrieval: 用于火星尺度地理空间检索的视觉-语言模型基准测试

Shuoyuan Wang, Yiran Wang, Hongxin Wei

机构 * Department of Statistics and Data Science(统计与数据科学系) Department of Earth and Space Sciences(地球与空间科学系)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 MarsRetrieval提出了一种用于评估视觉-语言模型在火星地理空间发现中检索能力的基准测试,强调领域特定微调对可推广发现的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏