arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2509.25711 2025-10-01 cs.CV 79%

ProbMed: A Probabilistic Framework for Medical Multimodal Binding

Yuan Gao, Sangwook Kim, Jianzhong You, Chris McIntosh

机构 * Peter Munk Cardiac Centre(彼得·默克心脏中心) Ted Rogers Centre for Heart Research(泰德·罗杰斯心脏病研究中心) University Health Network(大学健康网络) Joint Department of Medical Imaging(联合医学影像部门) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26378 2025-10-01 cs.IR cs.CV 79%

MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval

Junjie Zhou, Ze Liu, Lei Xiong, Jin-Ge Yao, Yueze Wang, Shitao Xiao, Fenfen Lin, Miguel Hu Chen, Zhicheng Dou, Siqi Bao, Defu Lian, Yongping Xiong, Zheng Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21151 2025-09-26 cs.CL cs.IR 79%

Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation Extraction

Lei Hei, Tingjing Liao, Yingxin Pei, Yiyang Qi, Jiaqi Wang, Ruiting Li, Feiliang Ren

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China(计算机科学与工程学院,东北大学,沈阳110819,中国)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20501 2025-09-26 cs.LG cs.CV 79%

Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules

Kishor Datta Gupta, Mohd Ariful Haque, Marufa Kamal, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Roy George

机构 * Clark Atlanta University(克拉克阿特兰大学) BRAC University(布拉克大学) United International University(国际联合大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15141 2025-09-26 cs.AI cs.LG physics.chem-ph 79%

Text-Augmented Multimodal LLMs for Chemical Reaction Condition Recommendation

Yu Zhang, Ruijie Yu, Kaipeng Zeng, Ding Li, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-罗克琴克研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19965 2025-09-25 cs.CV 79%

SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding

Phyo Thet Yee, Dimitrios Kollias, Sudeepta Mishra, Abhinav Dhall

机构 * IIT Ropar(印度IIT罗帕尔) Queen Mary University of London(伦敦女王玛丽大学) Monash University(墨尔本大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2026, project page : https://novicemm.github.io/synchrorama

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 79%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18284 2025-09-24 cs.CV 79%

Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction

Yi Gu, Kuniaki Saito, Jiaxin Ma

机构 * OMRON SINIC X Corporation(OMRON SINIC X公司) Nara Institute of Science and Technology(名取科学技術大學院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17762 2025-09-23 cs.CV 79%

Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction

Sitian Shen, Georgi Pramatarov, Yifu Tao, Daniele De Martini

机构 * Mobile Robotics Group (MRG), Oxford Robotics Institute, Department of Engineering Science, University of Oxford, UK(移动机器人组(MRG),牛津机器人研究所,工程科学系,牛津大学,英国)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15269 2025-09-23 cs.CV cs.LG 79%

Test-Time Multimodal Backdoor Detection by Contrastive Prompting

Yuwei Niu, Shuo He, Qi Wei, Zongyu Wu, Feng Liu, Lei Feng

机构 * Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) Penn State University(宾夕法尼亚州立大学) University of Melbourne(墨尔本大学) Southeast University(东南大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15211 2025-09-19 cs.CL 79%

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

Petros Stylianos Giouroukis, Dimitris Dimitriadis, Dimitrios Papadopoulos, Zhenwen Shao, Grigorios Tsoumakas

机构 * Aristotle University of Thessaloniki(亚里士多德大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12590 2025-09-19 cs.CV cs.LG 79%

Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering

Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck, Estelle Aflalo, Shao-Yen Tseng, Vasudev Lal, Phillip Howard

机构 * Oracle Intel Labs(英特尔实验室) Thoughtworks

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 Figures, 8 Tables. arXiv admin note: text overlap with arXiv:2410.13976

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13474 2025-09-18 cs.CV 79%

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Yujia Lin, Nicholas Evans

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(巴尔迪马十一点大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10282 2025-09-15 cs.CV cs.LG 79%

MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection

Gang Li, Tianjiao Chen, Mingle Zhou, Min Li, Delong Han, Jin Wan

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan)(计算能力网络与信息安全部教育部重点实验室,山东计算机科学中心(济南国家超级计算机中心)) Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算能力互联网与服务计算重点实验室,山东省计算机科学基础研究中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Page 14, 5 pictures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07666 2025-09-10 cs.CL cs.IR 79%

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval

Xixi Wu, Yanchao Tan, Nan Hou, Ruiyang Zhang, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) Fuzhou University(福州大学) University of Macau(澳门大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments EMNLP Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02017 2025-09-03 cs.IR cs.AI 79%

Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs

Yuhao Wang, Junwei Pan, Xinhang Li, Maolin Wang, Yuan Wang, Yue Liu, Dapeng Liu, Jie Jiang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments CIKM 2025 Full Research Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00751 2025-09-03 cs.CV 79%

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南国家大学科学学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21595 2025-09-03 cs.CV 79%

PS-ReID: Advancing Person Re-Identification and Precise Segmentation with Multimodal Retrieval

Jincheng Yan, Yun Wang, Xiaoyan Luo, Yu-Wing Tai

机构 * School of Astronautics, Beihang University(北京航空航天大学航天学院) Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07460 2025-08-28 cs.LG cs.AI cs.DB 79%

HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 79%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17044 2025-08-26 cs.CV cs.RO 79%

M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments

Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 29 pages, 3 figures, 13 tables. Preprint of the accepted article in Optical Memory and Neural Network Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01273 2025-08-26 cs.IR cs.MM 79%

SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries

Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments accepted to CBMI 2024 as a regular paper; https://github.com/simula/soccer-rag

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16882 2025-08-26 eess.IV cs.CV 79%

Multimodal Medical Endoscopic Image Analysis via Progressive Disentangle-aware Contrastive Learning

Junhao Wu, Yun Li, Junhao Li, Jingliang Bian, Xiaomao Fan, Wenbin Lei, Ruxin Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) First Affiliated Hospital, Sun Yat-sen University(中山大学第一附属医院) College of Big Data and Internet, Shenzhen Technology University(深圳技术大学大数据与互联网学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages,6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10264 2025-08-25 cs.LG cs.AI cs.IR 79%

Order-Preserving Dimension Reduction for Multimodal Semantic Embedding

Chengyu Gong, Gefei Shen, Luanzheng Guo, Nathan Tallent, Dongfang Zhao

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14515 2025-08-21 cs.IR cs.AI 79%

MISS: Multi-Modal Tree Indexing and Searching with Lifelong Sequential Behavior for Retrieval Recommendation

Chengcheng Guo, Junda She, Kuo Cai, Shiyao Wang, Qigen Hu, Qiang Luo, Kun Gai, Guorui Zhou

机构 * Kuaishou Inc.(快手公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22884 2025-08-20 cs.CV 79%

AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMs

Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya

机构 * University of Maryland, College Park(马里兰大学学院市分校) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12149 2025-08-19 cs.AI 79%

MOVER: Multimodal Optimal Transport with Volume-based Embedding Regularization

Haochen You, Baojing Liu

机构 * Columbia University(哥伦比亚大学) Hebei Institute of Communications(河北通信学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted as a conference paper at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17297 2025-08-08 cs.AI 79%

Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts

Zhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang, Xiaoyuan Yi, Yukun Yan, Ge Yu, Maosong Sun

机构 * Northeastern University, China(东北大学) Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19264 2025-08-07 cs.CV 79%

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

Sijie Li, Chen Chen, Jungong Han

机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03494 2025-08-06 cs.CV 79%

Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval

Shreyank N Gowda, Xiaobo Jin, Christian Wagner

机构 * School of Computer Science, The University of Nottingham, NG8 1BB Nottingham, U.K.(计算机科学学院,诺丁汉大学) Department of Intelligent Science, Xi’an Jiaotong-Liverpool University, China, 215123.(智能科学系,西安交通大学利物浦大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏