arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3437 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3437 篇

2601.04127 2026-01-08 cs.CV cs.AI 81%

Pixel-Wise Multimodal Contrastive Learning for Remote Sensing Images

像素级多模态对比学习用于遥感图像

Leandro Stival, Ricardo da Silva Torres, Helio Pedrini

机构 * Wageningen University & Research(瓦赫宁根大学与研究学院) Institute of Computing, University of Campinas (UNICAMP)(计算学院,坎皮纳斯大学(UNICAMP))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出像素级多模态对比学习方法,通过二维表示和对比学习提升遥感图像中特征提取效果,优于现有方法。

Comments 21 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25270 2026-01-06 cs.LG cs.AI cs.CV 81%

InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions

InfMasking:通过对比多模态交互释放协同信息

Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai, Dongkai Wang, Zhao Kang, Jun Wang, Zenglin Xu, Jiang Duan

机构 * School of Computing and Artificial Intelligence(计算与人工智能学院) Engineering Research Center of Intelligent Finance Ministry of Education(教育部长智能金融工程研究中心) Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences(中国科学院深圳先进技术研究院) University of Electronic Science and Technology of China(电子科技大学) Shanghai Academy of AI for Science(上海人工智能科学研究院) Artificial Intelligence Innovation and Incubation Institute Fudan University(复旦大学人工智能创新与孵化院) Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province(四川省人工智能与数字金融重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 InfMasking通过无限遮蔽策略增强多模态协同信息提取,实现多模态表示学习中的协同信息优化与性能提升。

Comments Conference on Neural Information Processing Systems (NeurIPS) 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18987 2025-12-23 cs.RO cs.CL cs.CV 81%

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

语义可感知的多模态检索:基于具身记忆的层次化移动操作

Ryosuke Korekata, Quanting Xie, Yonatan Bisk, Komei Sugiura

机构 * Keio University(keio大学) Keio AI Research Center(keio人工智能研究中心) Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本研究提出Affordance RAG框架,通过构建具有可操作性的具身记忆,提升机器人在开放词汇移动操作中的检索性能和任务成功率。

Comments Accepted to IEEE RA-L, with presentation at ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12935 2025-12-16 cs.CV cs.AI cs.IR 81%

Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion

通过级联嵌入-重排序和时间感知得分融合实现统一的交互式多模态时刻检索

Toan Le Ngo Thanh, Phat Ha Huu, Tan Nguyen Dang Duy, Thong Nguyen Le Minh, Anh Nguyen Nhu Tinh

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过级联嵌入-重排序和时间感知得分融合实现统一的多模态时刻检索系统,解决跨模态噪声、时间序列连贯性和模态选择问题。

Comments Accepted at AAAI Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02454 2025-12-16 cs.CL cs.AI 81%

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

多模态深度研究员:从零开始生成文本-图表交织报告的代理框架

Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang, Xingyu Liu, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu, Bo Zhang, Wei Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 多模态深度研究员通过代理框架实现从零开始生成文本-图表交织报告,利用FDV结构化文本表示提升可视化生成质量。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04741 2025-12-02 cs.AI cs.CL cs.HC 81%

Question Answering for Decisionmaking in Green Building Design: A Multimodal Data Reasoning Method Driven by Large Language Models

绿色建筑设计中的决策制定问答:一种由大语言模型驱动的多模态数据推理方法

Yihui Li, Xiaoyue Yan, Hao Zhou, Borong Lin

机构 * School of Architecture(建筑学院) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出GreenQA,一种由大语言模型驱动的多模态数据推理方法,用于提升绿色建筑设计决策效率。

Comments Published at Association for Computer Aided Design in Architecture (ACADIA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09878 2025-11-24 cs.CV cs.AI cs.RO 81%

CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation

CleverDistiller: 简单且空间一致的跨模态知识蒸馏

Hariprasath Govindarajan, Maciej K. Wozniak, Marvin Klingner, Camille Maurice, B Ravi Kiran, Senthil Yogamani

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 CleverDistiller通过简单有效的跨模态知识蒸馏方法,在自动驾驶任务中实现语义分割和3D目标检测的高性能表现。

Comments Accepted to BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19579 2025-11-20 cs.CV cs.AI cs.LG 81%

Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration

Francisco Mena, Dino Ienco, Cassio F. Dantas, Roberto Interdonato, Andreas Dengel

机构 * Department of Computer Science, University of Kaiserslautern-Landau (RPTU)(科斯拉尔特伦大学计算机科学系) SDS, German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) INRAE, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) CIRAD, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) INRIA, EVERGREEN, University of Montpellier(蒙彼利埃大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at the Machine Learning journal, CfP: Discovery Science 2024

Journal ref Machine Learning 114, 279 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11305 2025-11-19 cs.IR cs.AI cs.CV cs.LG 81%

MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising

Chenghan Fu, Daoze Zhang, Yukang Lin, Zhanheng Nie, Xiang Zhang, Jianyu Liu, Yueran Liu, Wanxian Guan, Pengjie Wang, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Taobao(淘宝)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05571 2025-11-11 cs.CV cs.AI 81%

C3-Diff: Super-resolving Spatial Transcriptomics via Cross-modal Cross-content Contrastive Diffusion Modelling

Xiaofei Wang, Stephen Price, Chao Li

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05404 2025-11-10 cs.CV cs.AI 81%

Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments

Laura Alejandra Encinar Gonzalez, John Folkesson, Rudolph Triebel, Riccardo Giubilato

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22900 2025-11-10 cs.CV cs.CL 81%

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

Mai A. Shaaban, Tausifa Jan Saleem, Vijay Ram Papineni, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Department of Mathematics and Computer Science, Faculty of Science, Alexandria University(亚历山大大学数学与计算机科学系) Sheikh Shakhbout Medical City(谢赫·沙赫布OUT医疗城)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22694 2025-10-28 cs.CV cs.CL cs.IR 81%

Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation

Shu Zhao, Tianyi Shen, Nilesh Ahuja, Omesh Tickoo, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025 UniReps Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20393 2025-10-24 cs.CV cs.MM 81%

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 81%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14605 2025-10-21 cs.CV cs.AI 81%

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14824 2025-10-17 cs.CL cs.CV cs.IR 81%

Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking

Ziqi Dai, Xin Zhang, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03477 2025-09-26 cs.LG cs.AI cs.CV 81%

Robult: Leveraging Redundancy and Modality Specific Features for Robust Multimodal Learning

Duy A. Nguyen, Abhi Kamboj, Minh N. Do

机构 * Siebel School of Computing and Data Science, UIUC, US(UIUC计算机与数据科学学院) Department of Electrical and Computer Engineering, UIUC, US(UIUC电气与计算机工程学院) VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学-伊利诺伊智能健康中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted and presented at IJCAI 2025 in Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21105 2025-09-23 cs.IR cs.AI cs.CL 81%

AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis

Callie C. Liao, Duoduo Liao, Sai Surya Gadiraju

机构 * Stanford University(斯坦福大学) George Mason University(乔治·马歇尔大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15882 2025-09-22 cs.CV cs.AI 81%

Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration

Xingmei Wang, Xiaoyu Hu, Chengkai Huang, Ziyan Zeng, Guohao Nie, Quan Z. Sheng, Lina Yao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15470 2025-09-22 cs.CV cs.AI 81%

Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

Thomas Z. Li, Aravind R. Krishnan, Lianrui Zuo, John M. Still, Kim L. Sandler, Fabien Maldonado, Thomas A. Lasko, Bennett A. Landman

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08338 2025-09-11 cs.CV cs.AI cs.LG 81%

Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

Jihyun Moon, Charmgil Hong

机构 * Handong Global University(-handong全球大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (MICCAI ISIC) 2025; 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01341 2025-09-03 cs.CV cs.AI 81%

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

Yunus Serhat Bicakci, Joseph Shingleton, Anahid Basiri

机构 * Vocational School of Social Sciences, Marmara University(马尔马拉大学社会科学职业学校) Geospatial Data Science Group, School of Geographical & Earth Sciences, University of Glasgow(格拉斯哥大学地理与地球科学学院空间数据科学小组)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00353 2025-09-03 cs.CV cs.AI 81%

AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data

Koushik Ahmed Kushal, Abdullah Al Mamun

机构 * Department of Computer Science Clarkson University(计算机科学系克拉克逊大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10435 2025-09-03 cs.CV cs.AI 81%

RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition

Xudong Yang, Yizhang Zhu, Hanfeng Liu, Zeyi Wen, Nan Tang, Yuyu Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19319 2025-08-28 eess.IV cs.AI cs.CV 81%

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

Pardis Moradbeiki, Nasser Ghadiri, Sayed Jalal Zahabi, Uffe Kock Wiil, Kristoffer Kittelmann Brockhattingen, Ali Ebrahimi

机构 * Department of Electrical and Computer Engineering, Isfahan University of Technology(电气与计算机工程系,伊斯法罕技术大学) SDU Health Informatics and Technology, The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(南部丹麦大学健康信息学与技术,马士基麦金尼莫勒研究所) Geriatric Research Unit, Department of Clinical Research, University of Southern Denmark(老年医学研究单元,临床研究系,南部丹麦大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14189 2025-08-15 cs.CL cs.AI 81%

DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base

Song Mao, Lejun Cheng, Pinlong Cai, Guohang Yan, Ding Wang, Botian Shi

机构 * Shanghai AI LAB(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments work in process

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06071 2025-08-15 cs.CV cs.MM 81%

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

Chang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja, Xiaokang Yang

机构 * Shanghai Jiao Tong University(上海交通大学) Singapore Institute of Technology(新加坡科技学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09170 2025-08-14 cs.LG cs.AI cs.CV cs.IR 81%

Multimodal RAG Enhanced Visual Description

Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz

机构 * Indian Institute of Technology (BHU)(印度理工学院(BHU)) University of Southampton(南安普顿大学) Modul University Vienna(维也纳应用科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM CIKM 2025. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14700 2025-08-14 cs.CV cs.AI 81%

Depth-Guided Self-Supervised Human Keypoint Detection via Cross-Modal Distillation

Aman Anand, Elyas Rashno, Amir Eskandari, Farhana Zulkernine

机构 * Queen’s University Ontario Canada(皇后大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏