arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2209.05463 2025-12-08 cs.CY cs.AI 79%

Modelling Business Agreements in the Multimodal Transportation Domain through Ontological Smart Contracts

通过本体智能合约建模多模态交通运输领域的商业协议

Mario Scrocca, Marco Comerio, Alessio Carenini, Irene Celino

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通过本体智能合约建模多模式交通运输领域的商业协议,展示其在拼车场景中的应用及优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05036 2025-12-03 cs.CL 79%

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

从词向量到多模态嵌入:大型语言模型的技术、应用与未来方向

Charles Zhang, Benji Peng, Xintian Sun, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Yichao Zhang, Xinyuan Song, Cheng Fei, Caitlyn Heqi Yin, Lawrence KQ Yan, Hongyang He, Tianyang Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Simon Fraser University(西蒙弗雷泽大学) Kyoto University(京都大学) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校) Emory University(埃默里大学) Cornell University(康奈尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The Hong Kong University of Science(香港科学大学) University of Liverpool(利物浦大学) University of Warwick(沃里克大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文综述了从词向量到多模态嵌入的发展,探讨了大型语言模型的技术、应用及未来方向。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22997 2025-12-01 cs.CV cs.RO 79%

MrGS: Multi-modal Radiance Fields with 3D Gaussian Splatting for RGB-Thermal Novel View Synthesis

MrGS: 多模态辐射场与3D高斯点云融合用于RGB-热成像新视角合成

Minseong Kweon, Janghyun Kim, Ukcheol Shin, Jinsun Park

机构 * Minnesota Robotics Institute (MnRI), University of Minnesota, Twin Cities(明尼苏达大学罗学院(MnRI)、明尼苏达大学双城分校) Department of Information Convergence Engineering (Artificial Intelligence Major), Pusan National University(信息融合工程系(人工智能专业),釜山国立大学) Department of Energy Engineering, Korea Institute of Energy Technology (KENTECH)(能源工程系,韩国能源技术研究所(KENTECH)) School of Computer Science and Engineering, Pusan National University(计算机科学与工程学院,釜山国立大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 MrGS通过多模态辐射场与3D高斯点云融合,实现RGB和热成像新视角合成,利用物理原理建模热传导和辐射现象,提升重建精度与效率。

Comments Accepted at Thermal Infrared in Robotics (TIRO) Workshop, ICRA 2025 (Best Poster Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19380 2025-11-25 cs.CV 79%

UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval

UISearch: 基于图的多模态企业UI截图检索

Maroun Ayli, Youssef Bakouny, Tushar Sharma, Nader Jalloul, Hani Seifeddine, Rima Kilany

机构 * Center For Computer Science(计算机科学中心) Saint Joseph University of Beirut(贝鲁特圣约瑟夫大学) Faculty of Computer Science(计算机科学学院) Dalhousie University(达尔豪斯大学) Murex

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 UISearch通过基于图的结构嵌入与语义检索结合,实现多模态企业UI截图检索,提升检索准确率与效率。

Comments 12 pages, 2 figures, 3 algorithms, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV 79%

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07419 2025-11-24 cs.IR cs.MM 79%

Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration

突破知识诅咒:通过知识软整合实现有效的多模态推荐

Kai Ouyang, Chen Tang, Zenghao Chai, Wenhao Zheng, Xiangjin Xie, Xuanji Xiao, Zhi Wang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出KSI框架,通过知识软整合解决多模态推荐中的知识诅咒问题,提升推荐个性化效果。

Comments Accepted to IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13189 2025-11-18 cs.CV cs.IR 79%

Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework

Diego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya, David Jiménez, Juan C. SanMiguel

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08181 2025-11-18 cs.IR cs.AI 79%

MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System

Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park

机构 * Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学) School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17714 2025-11-11 cs.CV 79%

F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model

Hanbo Bi, Zhiqiang Yuan, Zexi Jia, Jiapei Zhang, Chongyang Li, Peixiang Luo, Ying Deng, Xiaoyue Duan, Jinchao Zhang

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17494 2025-11-11 eess.IV cs.CV 79%

Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning

Jun-En Ding, Chien-Chin Hsu, Chi-Hsiang Chu, Shuqiang Wang, Feng Liu

机构 * Department of Systems Engineering, Stevens Institute of Technology, Hoboken, New Jersey, USA(系统工程系,史蒂文斯理工学院) Department of Nuclear Medicine, Kaohsiung Chang Gung Memorial Hospital, Kaohsiung, Taiwan(高雄长庚纪念医院核医学部) Institute of Statistics, National University of Kaohsiung, Kaohsiung, Taiwan(国立高雄大学统计研究所) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China(深圳先进技术研究院,中国科学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03070 2025-11-11 cs.LG cs.MM 79%

FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization

Manh Duong Nguyen, Trung Thanh Nguyen, Huy Hieu Pham, Trong Nghia Hoang, Phi Le Nguyen, Thanh Trung Huynh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Nagoya University(名古屋大学) Washington State University(华盛顿州立大学) Swiss Federal Institute of Technology Lausanne(洛桑联邦理工学院)

专题命中 跨模态检索 :cross-modal(title);multi-modal(abstract);分类 cs.MM

Comments The 22nd International Symposium on Network Computing and Applications (NCA 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06268 2025-11-11 cs.CV cs.CY 79%

LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

Jian Zhang, Junyi Guo, Junyi Yuan, Huanda Lu, Yanlin Zhou, Fangyu Wu, Qiufeng Wang, Dongming Lu

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Dunhuang Academy(敦煌研究院) Zhejiang University(浙江大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01892 2025-11-05 cs.LG cs.CL 79%

Retrieval-Augmented Multimodal Depression Detection

Ruibo Hou, Shiyu Teng, Jiaqing Liu, Shurong Chai, Yinhao Li, Lanfen Lin, Yen-Wei Chen

机构 * College of Information Science and Engineering, Ritsumeikan University(信息科学与工程学院,立命馆大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted in IEEE EMBC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00903 2025-11-04 cs.CL 79%

ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval

Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan, Rabiul Awal, Shambhavi Mishra, Akshay Kalkunte Suresh, Srivatsava Daruru, Enamul Hoque, Spandana Gella, Torsten Scholak, Sai Rajeswar

机构 * ServiceNow York University(约克大学) MILA - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) École de technologie supérieure(卓越工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27350 2025-11-03 cs.CV 79%

RzenEmbed: Towards Comprehensive Multimodal Retrieval

Weijian Jian, Yajun Zhang, Dawei Liang, Chunyu Xie, Yixiao He, Dawei Leng, Yuhui Yin

机构 * AI Research(360人工智能研究院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23224 2025-10-28 cs.CV cs.IR 79%

Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

Hongyi Wang, Zhengjie Zhu, Jiabo Ma, Fang Wang, Yue Shi, Bo Luo, Jili Wang, Qiuyu Cai, Xiuming Zhang, Yen-Wei Chen, Lanfen Lin, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Department of Radiology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属同济医院放射科) Department of Pathology, Sir Run Run Shaw Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属邵氏医院病理科) Department of Pathology, The Central Hospital of Wuhan, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属武汉中心医院病理科) Department of Pathology, The First Affiliated Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属第一医院病理科) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科学与技术大学生命科学系)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22880 2025-10-28 cs.LG cs.AI 79%

Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data

Duong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Phi Le Nguyen

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Washington State University(华盛顿州立大学) VinUniversity(文大学) Griffin University(格里芬大学) Hanoi University of Science and Technology(河内科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05715 2025-10-28 cs.IR cs.MM 79%

From ID-based to ID-free: Rethinking ID Effectiveness in Multimodal Collaborative Filtering Recommendation

Guohao Li, Li Jing, Jia Wu, Xuefei Li, Kai Zhu, Yue He

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments We identified that our current approach achieves its reported performance only under specific data conditions, and its robustness is weaker than we initially expected

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18303 2025-10-22 cs.CV 79%

Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models

Lehan Wang, Yi Qin, Honglong Yang, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15547 2025-10-20 cs.AI cs.ET cs.LG cs.SY eess.SP eess.SY 79%

Hypergraph Contrastive Sensor Fusion for Multimodal Fault Diagnosis in Induction Motors

Usman Ali, Ali Zia, Waqas Ali, Umer Ramzan, Abdul Rehman, Muhammad Tayyab Chaudhry, Wei Xiang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to IEEE Sensors Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09585 2025-10-20 cs.CV 79%

Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation

Jitesh Jain, Zhengyuan Yang, Humphrey Shi, Jianfeng Gao, Jianwei Yang

机构 * Microsoft Research, Redmond(微软研究院(红mond)) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 跨模态检索 :multimodal(title);MLLM(abstract);分类 cs.CV

Comments Project Page: https://praeclarumjj3.github.io/visper_lm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14136 2025-10-17 cs.AI 79%

A Multimodal Approach to Heritage Preservation in the Context of Climate Change

David Roqui, Adèle Cormier, nistor Grozavu, Ann Bourges

机构 * ETIS laboratory, Cergy university(Cergy大学ETIS实验室) Research and Restoration Center for the Museums of France (C2RMF)(法国博物馆研究与修复中心(C2RMF))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12801 2025-10-15 cs.CV cs.IR 79%

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Kartik Narayan, Yang Xu, Tian Cao, Kavya Nerella, Vishal M. Patel, Navid Shiee, Peter Grasch, Chao Jia, Yinfei Yang, Zhe Gan

机构 * Johns Hopkins University(约翰霍普金斯大学) Apple(苹果公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10828 2025-10-14 cs.IR cs.AI 79%

VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering

Zhenghan Tai, Hanwei Wu, Qingchen Hu, Jijun Chi, Hailin He, Lei Ding, Tung Sum Thomas Kwok, Bohuai Xiao, Yuchen Hua, Suyuchen Wang, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Jerry Huang, Jiayi Zhang, Gonghao Zhang, Chaolong Jiang, Jingrui Tian, Sicheng Lyu, Zeyu Li, Boyu Han, Fengran Mo, Xinyue Yu, Yufei Cui, Ling Zhou, Xinyu Wang

机构 * University of Toronto(多伦多大学) McMaster University(麦马斯特大学) McGill University(麦吉尔大学) University of Manitoba(曼尼托巴大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Montreal(蒙特利尔大学) Mila CUHK(香港中文大学) HKUST(GZ)(香港理工大学(广州)) Nanyang Technological University(南洋理工大学) Stanford University(斯坦福大学) CG Matrix Technology Limited(CG矩阵科技有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04379 2025-10-13 cs.CV cs.LG 79%

VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones

Lefei Shen, Mouxiang Chen, Xu Liu, Han Fu, Xiaoxue Ren, Jianling Sun, Zhuo Li, Chenghao Liu

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) State Street Technology (Zhejiang) Ltd.(State Street Technology(浙江)有限公司) Salesforce Research Asia(Salesforce亚洲研究)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03268 2025-10-09 cs.LG cs.AI 79%

Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment

Lingjie Yi, Raphael Douady, Chao Chen

机构 * Stony Brook University(史坦尼·布鲁克大学) University Paris 1 Pantheon-Sorbonne(巴黎第一大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11465 2025-10-07 cs.CL cs.LG 79%

CEMTM: Contextual Embedding-based Multimodal Topic Modeling

Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00579 2025-10-06 cs.MM cs.IR 79%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02580 2025-10-06 cs.AI 79%

V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving

Xuewen Luo, Fengze Yang, Fan Ding, Xiangbo Gao, Shuo Xing, Yang Zhou, Zhengzhong Tu, Chenxi Liu

机构 * University of Utah(犹他大学) Monash University(莫纳什大学) Texas A&M University(德克萨斯农工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06461 2025-10-06 cs.CV 79%

Ranked from Within: Ranking Large Multimodal Models Without Labels

Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

机构 * Australian National University Sydney AI Centre, The University of Sydney Curtin University University of \'OBuda

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏