arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2507.04870 2025-11-17 cs.LG 78%

NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification

Jun Hu, Yufei He, Yuan Li, Bryan Hooi, Bingsheng He

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08970 2025-11-13 astro-ph.SR 78%

JW-Flare: Accurate Solar Flare Forecasting Method Based on Multimodal Large Language Models

Mingfu Shao, Hui Wang, Yuyang Li, Jiaben Lin, Jifeng Liu, Baolin Tan, Juan Guo, Yin Zhang, Jing Huang, Jiangtao Su, Yingzi Sun, Haiqing Xu, Jie Chen, Suo Liu, Yuanyong Deng, Liyue Tong, Yang Bai, Cunshi Wang, Kaifan Ji, Yuqing Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06723 2025-11-11 cs.LG 78%

Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation

Evelyn Chee, Wynne Hsu, Mong Li Lee

机构 * School of Computing, National University of Singapore(computing学院,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Accepted to ECAI 2025

Journal ref 28th European Conference on Artificial Intelligence (ECAI), 2025, pp.1083-1090

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05716 2025-11-11 cs.LG 78%

Distributionally Robust Multimodal Machine Learning

Peilin Yang, Yu Ma

机构 * University of Wisconsin, Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03196 2025-11-06 cs.LG stat.ML 78%

Cross-Modal Alignment via Variational Copula Modelling

Feng Wu, Tsai Hor Chan, Fuying Wang, Guosheng Yin, Lequan Yu

机构 * School of Computing and Data Science, University of Hong Kong, Hong Kong, China(计算与数据科学学院,香港大学,香港,中国)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract)

Journal ref published by ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00949 2025-11-04 cs.LG 78%

Motion-Robust Multimodal Fusion of PPG and Accelerometer Signals for Three-Class Heart Rhythm Classification

Yangyang Zhao, Matti Kaisti, Olli Lahdenoja, Tero Koivisto

机构 * Department of Computing, Faculty of Technology, University of Turku(图波大学计算系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted for publication in the Companion of the 2025 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2025 International Symposium on Wearable Computers (UbiComp/ISWC 2025 Companion). 5 pages, 3 figures. Author's accepted manuscript (AAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21551 2025-10-27 cs.LG 78%

Interpretable Multimodal Zero-Shot ECG Diagnosis via Structured Clinical Knowledge Alignment

Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed

机构 * Eindhoven University of Technology(埃因霍温理工大学) Singapore Management University(新加坡管理大学) Maxima Medical Center(马克斯医疗中心) Erasmus Medical Center(埃因霍温医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20976 2025-10-27 cs.LG 78%

L^2M^3OF: A Large Language Multimodal Model for Metal-Organic Frameworks

Jiyu Cui, Fang Wu, Haokai Zhao, Minggao Feng, Xenophon Evangelopoulos, Andrew I. Cooper, Yejin Choi

机构 * Department of Chemistry, University of Liverpool(利兹大学化学系) Leverhulme Research Centre for Functional Materials Design, University of Liverpool(利兹大学功能性材料设计研究所以) Department of Computer Science, University of Stanford(斯坦福大学计算机科学系) School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18240 2025-10-22 cs.LG 78%

Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment

Haobin Li, Yijie Lin, Peng Hu, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) National Key Laboratory of Fundamental Algorithms and Models for Engineering Numerical Simulation, Sichuan University(四川省工程数值模拟基础算法与模型国家重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 30 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16086 2025-10-21 cs.LG stat.AP 78%

FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment Analysis

Ziyang Liu, Pengjunfei Chu, Shuming Dong, Chen Zhang, Mingcheng Li, Jin Wang

机构 * School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院) School of Advanced Manufacturing Engineering, Hefei University(合肥大学先进制造工程学院) College of Global Talents, BITZH, Beijing Institute of Technology(北京理工大学珠海学院全球人才学院) Academy for Engineering and Technology, Fudan University(复旦大学工程与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages,3 figures

Journal ref In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15953 2025-10-21 cs.CR 78%

Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations

Sisir Doppalapudi

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01719 2025-10-21 cs.LG 78%

Robust Anomaly Detection through Multi-Modal Autoencoder Fusion for Small Vehicle Damage Detection

Sara Khan, Mehmed Yüksel, Frank Kirchner

机构 * Faculty of Mathematics and Computer Science, University of Bremen(数学与计算机科学学院,不莱梅大学) Robotics Innovation Center, Deutsches Forschungszentrum für Künstliche Intelligenz(机器人创新中心,德国人工智能研究中心) Engineering Software Communication, Robert Bosch GmbH(工程软件通信,罗伯特·博世有限公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 17 pages, 12 figures, submitted to Elsevier MLWA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04863 2025-10-21 cs.LG q-bio.BM 78%

OneProt: Towards Multi-Modal Protein Foundation Models

Klemens Flöge, Srisruthi Udayakumar, Johanna Sommer, Marie Piraud, Stefan Kesselheim, Vincent Fortuin, Stephan Günneman, Karel J van der Weg, Holger Gohlke, Erinc Merdivan, Alina Bazarova

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments 34 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11666 2025-10-14 eess.SP cs.LG 78%

Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion

Timothy Oladunni, Ehimen Aneni

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22962 2025-10-07 cs.LG cond-mat.mtrl-sci physics.chem-ph 78%

Multimodal machine learning with large language embedding model for polymer property prediction

Tianren Zhang, Dai-Bei Yang

机构 * Department of Materials Science and Engineering, University of Delaware, Newark, Delaware 19716, United States(材料科学与工程系,德雷克塞尔大学) Department of Chemistry, University of Pennsylvania, Philadelphia, Pennsylvania 19104, United States(化学系,宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref Chem. Mater. 2025, 37, 7002-7013

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25525 2025-10-01 cs.CR cs.LG 78%

Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models

Boyang Zhang, Istemi Ekin Akkus, Ruichuan Chen, Alice Dethise, Klaus Satzke, Ivica Rimac, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24431 2025-09-30 cs.LG 78%

Semantic Compression via Multimodal Representation Learning

Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. of Information Engineering, Electronics, and Telecomm.(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20840 2025-09-26 cs.LG 78%

Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition

Jiaqi Tang, Yinsong Xu, Yang Liu, Qingchao Chen

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学人民医院医学技术研究所) National Institute of Health Data Science(国家健康数据科学研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18005 2025-09-23 cs.RO 78%

M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer

Yanxin Zhang, Liang He, Zeyi Kang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学 西安) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学 巴黎)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13636 2025-09-18 cs.LG 78%

Multimodal signal fusion for stress detection using deep neural networks: a novel approach for converting 1D signals to unified 2D images

Yasin Hasanpoor, Bahram Tarvirdizadeh, Khalil Alipour, Mohammad Ghamari

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 14 pages 7 images 2 tables

Journal ref 11760_2025_4734_Article

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05300 2025-09-17 eess.SP 78%

Harnessing Multimodal Sensing for Multi-user Beamforming in mmWave Systems

Kartik Patel, Robert W. Heath

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref IEEE Trans.Wireless Commun. 23 (2024) 18725-18739

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12930 2025-09-17 cs.DC 78%

Analysis and Optimization of Wireless Multimodal Federated Learning on Modal Heterogeneity

Xuefeng Han, Wen Chen, Jun Li, Ming Ding, Qingqing Wu, Kang Wei, Xiumei Deng, Yumeng Shao, Qiong Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10720 2025-09-17 cs.CY 78%

Adapting Public Personas: A Multimodal Study of U.S. Legislators' Cross-Platform Social Media Strategies

Weihong Qi, Anushka Dave, Chen Ling

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11187 2025-09-16 cs.CR 78%

DMLDroid: Deep Multimodal Fusion Framework for Android Malware Detection with Resilience to Code Obfuscation and Adversarial Perturbations

Doan Minh Trung, Tien Duc Anh Hao, Luong Hoang Minh, Nghi Hoang Khoa, Nguyen Tan Cam, Van-Hau Pham, Phan The Duy

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03837 2025-09-05 cs.LG cs.IT math.IT 78%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03029 2025-09-04 cs.LG 78%

Multimodal learning of melt pool dynamics in laser powder bed fusion

Satyajit Mojumder, Pallock Halder, Tiana Tonge

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 20 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02275 2025-09-03 cs.RO 78%

Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals

Fengyi Wang, Xiangyu Fu, Nitish Thakor, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学) Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18551 2025-08-27 cs.LG 78%

BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration

Jun Hou, Le Wang, Xuan Wang

机构 * Department of Computer Science, Virginia Tech, Blacksburg, VA, USA(计算机科学系,弗吉尼亚理工学院,布莱克斯堡,VA,美国) Department of Agricultural and Applied Economics, Virginia Tech, Blacksburg, VA, USA(农业与应用经济学系,弗吉尼亚理工学院,布莱克斯堡,VA,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17640 2025-08-26 eess.SP 78%

Multimodal Radio and Vision Fusion for Robust Localization in Urban V2I Communications

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Henk Wymeersch

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages, 6 figures, submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14485 2025-08-22 cs.IR 78%

Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion

Moyu Zhang, Yongxiang Tang, Yujun Jin, Jinxin Hu, Yu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025, 11 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏