arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2507.06071 2025-08-15 cs.CV cs.MM 81%

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

Chang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja, Xiaokang Yang

机构 * Shanghai Jiao Tong University(上海交通大学) Singapore Institute of Technology(新加坡科技学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09170 2025-08-14 cs.LG cs.AI cs.CV cs.IR 81%

Multimodal RAG Enhanced Visual Description

Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz

机构 * Indian Institute of Technology (BHU)(印度理工学院(BHU)) University of Southampton(南安普顿大学) Modul University Vienna(维也纳应用科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM CIKM 2025. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14700 2025-08-14 cs.CV cs.AI 81%

Depth-Guided Self-Supervised Human Keypoint Detection via Cross-Modal Distillation

Aman Anand, Elyas Rashno, Amir Eskandari, Farhana Zulkernine

机构 * Queen’s University Ontario Canada(皇后大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06154 2025-08-11 cs.IR cs.AI cs.MM 81%

Semantic Item Graph Enhancement for Multimodal Recommendation

Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Dusit Niyato, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14675 2025-07-22 cs.CV cs.CL 81%

Docopilot: Improving Multimodal Models for Document-Level Understanding

Yuchen Duan, Zhe Chen, Yusong Hu, Weiyun Wang, Shenglong Ye, Botian Shi, Lewei Lu, Qibin Hou, Tong Lu, Hongsheng Li, Jifeng Dai, Wenhai Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学) Nankai University(南开大学) Fudan University(复旦大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14485 2025-07-22 cs.CV cs.AI 81%

Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion

Hongye Hou, Liu Zhan, Yang Yang

机构 * Xi'an Jiaotong University(西安交通大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08683 2025-07-14 cs.CV cs.AI 81%

MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing

Debashis Gupta, Aditi Golder, Rongkhun Zhu, Kangning Cui, Wei Tang, Fan Yang, Ovidiu Csillik, Sarra Alaqahtani, V. Paul Pauca

机构 * City University of Hong Kong(香港城市大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09822 2025-07-09 cs.CV cs.AI 81%

Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model

Camille Delgrange, Olga Demler, Samia Mora, Bjoern Menze, Ezequiel de la Rosa, Neda Davoudi

机构 * Signal Processing Institute EPFL University Lausanne(瑞士洛桑联邦理工学院信号处理研究所) Brigham and Women’s Hospital Harvard Medical School(哈佛医学院布里奇沃特医院) Department of Quantitative Biomedicine University of Zurich(苏黎世大学定量生物医学系) ETH AI Center, Department of Computer Science Department of Quantitative Biomedicine University of Zurich(苏黎世大学定量生物医学系)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as oral paper at AIM-FM workshop, Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03868 2025-07-08 cs.AI cs.CE cs.CY cs.MM 81%

From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM

Xinyi Wu, Yanhao Jia, Luwei Xiao, Shuai Zhao, Fengkuang Chiang, Erik Cambria

机构 * School of Education, Shanghai Jiao Tong University(上海交通大学教育学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学 computing and Data Science 学院) school of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08541 2025-07-08 cs.CV cs.AI 81%

TrajFlow: Multi-modal Motion Prediction via Flow Matching

Qi Yan, Brian Zhang, Yutong Zhang, Daniel Yang, Joshua White, Di Chen, Jiachao Liu, Langechuan Liu, Binnan Zhuang, Shaoshuai Shi, Renjie Liao

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(向量人工智能研究所) University of Waterloo(滑铁卢大学) Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学) Tesla(特斯拉) XPeng Motors(小鹏汽车) Leapmotor Nvidia(英伟达) DiDi(滴滴出行) Canada CIFAR AI Chair(加拿大 CIFAR 人工智能主席)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00041 2025-07-02 cs.AI cs.CV cs.IR 81%

TalentMine: LLM-Based Extraction and Question-Answering from Multimodal Talent Tables

Varun Mannam, Fang Wang, Chaochun Liu, Xin Chen

机构 * Amazon(亚马逊)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Submitted to KDD conference, workshop: Talent and Management Computing (TMC 2025), https://tmcworkshop.github.io/2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17777 2025-07-01 cs.CV cs.AI 81%

Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification

Raja Kumar, Raghav Singhal, Pranamya Kulkarni, Deval Mehta, Kshitij Jadhav

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔) AIM for Health Lab, Department of Data Science & AI, Monash University(AIM健康实验室,数据科学与人工智能系,莫纳什大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Transactions on Machine Learning Research (TMLR). Raja Kumar and Raghav Singhal contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13056 2025-06-27 cs.AI cs.CV cs.LG 81%

Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning

Haibo Qiu, Xiaohan Lan, Fanfan Liu, Xiaohu Sun, Delian Ruan, Peng Shi, Lin Ma

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Project Page: https://github.com/MM-Thinking/Metis-RISE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24164 2025-06-18 cs.CL cs.CV 81%

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Shilin Xu, Yanwei Li, Rui Yang, Tao Zhang, Yueyi Sun, Wei Chow, Linfeng Li, Hang Song, Qi Xu, Yunhai Tong, Xiangtai Li, Hao Fei

机构 * ByteDance(字节跳动) Peking University(北京大学) National University of Singapore(新加坡国立大学)

专题命中 跨模态检索 :multimodal(title);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12232 2025-06-17 cs.CV cs.CL 81%

Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles

Mohammed Elhenawy, Shadi Jaradat, Taqwa I. Alhadidi, Huthaifa I. Ashqar, Ahmed Jaber, Andry Rakotonirainy, Mohammad Abu Tami

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01120 2025-06-17 cs.CV cs.AI 81%

Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning

Jian Lang, Zhangtao Cheng, Ting Zhong, Fan Zhou

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 8 figures. Accepted by AAAI 2025. Codes are released at https://github.com/Jian-Lang/RAGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11063 2025-06-16 cs.CL cs.AI 81%

Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation

Jiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Yuyao Ge, Zhecheng Li, Xueqi Cheng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学梅德福分校) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07296 2025-06-10 cs.IR cs.AI cs.CV 81%

HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval

Arian Askari, Emmanouil Stergiadis, Ilya Gusev, Moran Beladev

机构 * Leiden University(莱顿大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at ACL 2025, Main track. 13 Pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07050 2025-06-10 cs.CV cs.IR cs.MM 81%

From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion

Zheng Wang, Kai Ying, Bin Xu, Chunjiao Wang, Cong Bai

机构 * College of Computer Science, Zhejiang University of Technology(浙江工业大学计算机科学学院) National Meteorological Information Center(国家气象信息中心) College of Computer Science, Zhejiang University of Technology& Zhejiang Key Laboratory of Visual Information Intelligent Processing(浙江工业大学计算机科学学院&浙江视觉信息智能处理重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06144 2025-06-09 cs.CV cs.CL cs.IR 81%

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

David Wan, Han Wang, Elias Stengel-Eskin, Jaemin Cho, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 18 pages. Code and data: https://github.com/meetdavidwan/clamr

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19868 2025-06-09 cs.IR cs.AI cs.CV cs.LG 81%

GENIUS: A Generative Framework for Universal Multimodal Search

Sungyeon Kim, Xinliang Zhu, Xiaofan Lin, Muhammet Bastan, Douglas Gray, Suha Kwak

机构 * Amazon(亚马逊公司) POSTECH

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02544 2025-06-06 cs.CL cs.AI cs.IR 81%

CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG

Yang Tian, Fan Liu, Jingyuan Zhang, Victoria W., Yupeng Hu, Liqiang Nie

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13107 2025-05-28 cs.CV cs.AI 81%

ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

Hao Yin, Guangzong Si, Zilei Wang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23128 2025-05-26 cs.SD cs.AI eess.AS 81%

CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining

Tristan Tsoi, Jiajun Deng, Yaolong Ju, Benno Weck, Holger Kirchhoff, Simon Lui

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI、eess.AS

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15840 2025-05-21 cs.CV cs.AI 81%

Masked Contrastive Reconstruction for Cross-modal Medical Image-Report Retrieval

Zeqiang Wei, Kai Jin, Xiuzhuang Zhou

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Journal ref Under review at Pattern Recognition Letters, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15135 2025-04-22 cs.IR cs.AI cs.CL 81%

KGMEL: Knowledge Graph-Enhanced Multimodal Entity Linking

Juyeon Kim, Geon Lee, Taeuk Kim, Kijung Shin

机构 * KAIST(韩国科学技术院) Hanyang University(翰林大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments SIGIR 2025 (Short)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13172 2025-04-18 cs.IR cs.CL cs.MM 81%

SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs

Haoxuan Li, Yi Bin, Yunshan Ma, Guoqing Wang, Yang Yang, See-Kiong Ng, Tat-Seng Chua

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12330 2025-04-18 cs.CL cs.AI 81%

HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation

Pei Liu, Xin Liu, Ruoyu Yao, Junming Liu, Siyuan Meng, Ding Wang, Jun Ma

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08748 2025-04-15 cs.IR cs.AI cs.CL cs.ET cs.LG 81%

A Survey of Multimodal Retrieval-Augmented Generation

Lang Mei, Siyu Mo, Zhihan Yang, Chong Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19390 2025-04-15 eess.IV cs.AI cs.CV 81%

Multi-modal Contrastive Learning for Tumor-specific Missing Modality Synthesis

Minjoo Lim, Bogyeong Kang, Tae-Eui Kam

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏