arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2407.14058 2025-05-27 cs.LG 78%

Towards the Causal Complete Cause of Multi-Modal Representation Learning

Jingyao Wang, Siyu Zhao, Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Fuchun Sun, Hui Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04452 2025-05-23 cs.IR 78%

COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation

Jinfeng Xu, Zheyu Chen, Wei Wang, Xiping Hu, Sang-Wook Kim, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11262 2025-05-19 cs.NE 78%

A Step towards Interpretable Multimodal AI Models with MultiFIX

Mafalda Malafaia, Thalea Schlender, Tanja Alderliesten, Peter A. N. Bosman

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 9 pages, 6 figures, submitted to GECCO conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10003 2025-05-16 cs.LG eess.SP 78%

AI2MMUM: AI-AI Oriented Multi-Modal Universal Model Leveraging Telecom Domain Large Model

Tianyu Jiao, Zhuoran Xiao, Yihang Huang, Chenhui Ye, Yijia Feng, Liyu Cai, Jiang Chang, Fangkun Liu, Yin Xu, Dazhi He, Yunfeng Guan, Wenjun Zhang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学 cooperative medianet innovation center) Nokia Bell Labs(诺基亚贝尔实验室) Institute of Intelligent Communications and Network Security, Chongqing University of Posts and Telecommunications(重庆邮电大学智能通信与网络安全研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18024 2025-05-16 cs.LG 78%

Multimodal Learning with Uncertainty Quantification based on Discounted Belief Fusion

Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier

机构 * Aix-Marseille Univ(艾克斯-马赛大学) LIS(实验室) IRD(法国国家科研 Institute) ESPACE-DEV

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref Proceedings of The 28th International Conference on Artificial Intelligence and Statistics 2025, in Proceedings of Machine Learning Research 258:3142-3150 Available from https://proceedings.mlr.press/v258/bezirganyan25a.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04120 2025-05-14 cs.LG 78%

Transformer representation learning is necessary for dynamic multi-modal physiological data on small-cohort patients

Bingxu Wang, Min Ge, Kunzhi Cai, Yuqi Zhang, Zeyi Zhou, Wenjiao Li, Yachong Guo, Wei Wang, Qing Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05698 2025-05-12 cs.LG 78%

Unsupervised Multi-modal Feature Alignment for Time Series Representation Learning

Chen Liang, Donghua Yang, Zhiyu Liang, Hongzhi Wang, Zheng Liang, Xiyang Zhang, Jianfeng Huang

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04634 2025-05-09 cs.LG cs.CE 78%

MatMMFuse: Multi-Modal Fusion model for Material Property Prediction

Abhiroop Bhattacharya, Sylvain G. Cloutier

机构 * Department of Electrical Engineering, École de technologie supérieure, Montréal, Canada(电气工程系,超技术学院,蒙特利尔,加拿大)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Presented at AI for Accelerated Materials Design(AI4Mat), ICLR 2025 (https://openreview.net/forum?id=pN4Zg6HBlq#discussion)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02168 2025-05-06 cs.AR 78%

CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design

Wenji Fang, Shang Liu, Jing Wang, Zhiyao Xie

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by ICLR 2025 (https://openreview.net/forum?id=rbnf7oe6JQ)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01945 2025-05-06 cs.MA cs.RO 78%

Act Natural! Extending Naturalistic Projection to Multimodal Behavior Scenarios

Hamzah I. Khan, David Fridovich-Keil

机构 * Department of Aerospace Engineering and Engineering Mechanics , University of Texas at Austin(航空航天工程与工程力学系,德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01135 2025-05-05 cs.LG 78%

Dual-Forecaster: A Multimodal Time Series Model Integrating Descriptive and Predictive Texts

Wenfa Wu, Guanyu Zhang, Zheng Tan, Yi Wang, Hongsheng Qi

机构 * Lenovo Research(联想研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00842 2025-05-05 cs.RO cs.SY eess.SY math.GR 78%

Fault-Tolerant Multi-Modal Localization of Multi-Robots on Matrix Lie Groups

Mahboubeh Zarei, Robin Chhabra

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00176 2025-05-02 cs.CE 78%

Generative Multimodal Multiscale Data Fusion for Digital Twins in Aerosol Jet Electronics Printing

Fatemeh Elhambakhsh, Suk Ki Lee, Hyunwoong Ko

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21826 2025-05-01 cs.RO 78%

An Underwater, Fault-Tolerant, Laser-Aided Robotic Multi-Modal Dense SLAM System for Continuous Underwater In-Situ Observation

Yaming Ou, Junfeng Fan, Chao Zhou, Pengju Zhang, Zongyuan Shen, Yichen Fu, Xiaoyan Liu, Zengguang Hou

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16524 2025-04-24 cs.IR 78%

Modality Reliability Guided Multimodal Recommendation

Xue Dong, Xuemeng Song, Na Zheng, Sicheng Zhao, Guiguang Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14927 2025-04-22 cs.HC 78%

Multimodal Non-Semantic Feature Fusion for Predicting Segment Access Frequency in Lecture Archives

Ruozhu Sheng, Jinghong Li, Shinobu Hasegawa

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 16 pages, 7 figures. Preliminary version; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14833 2025-04-22 cs.NI cs.CR 78%

IoT-AMLHP: Aligned Multimodal Learning of Header-Payload Representations for Resource-Efficient Malicious IoT Traffic Classification

Fengyuan Nie, Guangjie Liu, Weiwei Liu, Jianan Huang, Bo Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13465 2025-04-21 cs.LG 78%

Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation

Duy A. Nguyen, Quan Huu Do, Khoa D. Doan, Minh N. Do

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11490 2025-04-18 cs.LG stat.ME 78%

Interventional Imbalanced Multi-Modal Representation Learning via $β$-Generalization Front-Door Criterion

Yi Li, Fei Song, Changwen Zheng, Jiangmeng Li, Fuchun Sun, Hui Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12025 2025-04-17 cs.LG 78%

FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning

Yu Zhang, Qingfeng Du, Jiaqi Lv

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09260 2025-04-15 cs.AR cs.LG 78%

NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed Graph

Wenji Fang, Wenkai Li, Shang Liu, Yao Lu, Hongce Zhang, Zhiyao Xie

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by Design Automation Conference (DAC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12501 2025-04-03 cs.IR 78%

Improving Multi-modal Recommender Systems by Denoising and Aligning Multi-modal Content and User Feedback

Guipeng Xv, Xinyu Li, Ruobing Xie, Chen Lin, Chong Liu, Feng Xia, Zhanhui Kang, Leyu Lin

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments After further review, we believe the content of the paper is not yet fully ready and requires additional time for improvement. To ensure quality, we have decided to withdraw this preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21964 2025-03-31 cs.LG q-bio.NC 78%

NeuroLIP: Interpretable and Fair Cross-Modal Alignment of fMRI and Phenotypic Text

Yanting Yang, Xiaoxiao Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20823 2025-03-28 cs.CR 78%

Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy

Joonhyun Jeong, Seyun Bae, Yeonsung Jung, Jaeryong Hwang, Eunho Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19594 2025-03-26 cs.IT eess.SP math.IT 78%

Perception-Enhanced Multitask Multimodal Semantic Communication for UAV-Assisted Integrated Sensing and Communication System

Ziji Guo, Haonan Tong, Zhilong Zhang, Danpu Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref WS21 ICC 2025 Workshop - ISCLAN

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00723 2025-03-24 cs.LG 78%

Re-Imagining Multimodal Instruction Tuning: A Representation View

Yiyang Liu, James Chenhao Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail Dianat, Raghuveer Rao, Lifu Huang, Dongfang Liu, Qifan Wang, Cheng Han

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15312 2025-03-20 astro-ph.GA 78%

Euclid Quick Data Release (Q1) Exploring galaxy properties with a multi-modal foundation model

Euclid Collaboration, M. Siudek, M. Huertas-Company, M. Smith, G. Martinez-Solaeche, F. Lanusse, S. Ho, E. Angeloudi, P. A. C. Cunha, H. Domínguez Sánchez, M. Dunn, Y. Fu, P. Iglesias-Navarro, J. Junais, J. H. Knapen, B. Laloux, M. Mezcua, W. Roster, G. Stevens, J. Vega-Ferrero, N. Aghanim, B. Altieri, A. Amara, S. Andreon, N. Auricchio, H. Aussel, C. Baccigalupi, M. Baldi, S. Bardelli, P. Battaglia, A. Biviano, A. Bonchi, E. Branchini, M. Brescia, J. Brinchmann, S. Camera, G. Cañas-Herrera, V. Capobianco, C. Carbone, J. Carretero, S. Casas, F. J. Castander, M. Castellano, G. Castignani, S. Cavuoti, K. C. Chambers, A. Cimatti, C. Colodro-Conde, G. Congedo, C. J. Conselice, L. Conversi, Y. Copin, F. Courbin, H. M. Courtois, M. Cropper, A. Da Silva, H. Degaudenzi, G. De Lucia, A. M. Di Giorgio, J. Dinis, C. Dolding, H. Dole, F. Dubath, C. A. J. Duncan, X. Dupac, S. Dusini, S. Escoffier, M. Farina, R. Farinelli, F. Faustini, S. Ferriol, F. Finelli, S. Fotopoulou, M. Frailis, E. Franceschi, S. Galeotta, K. George, B. Gillis, C. Giocoli, J. Gracia-Carpio, B. R. Granett, A. Grazian, F. Grupp, S. Gwyn, S. V. H. Haugan, W. Holmes, I. M. Hook, F. Hormuth, A. Hornstrup, K. Jahnke, M. Jhabvala, E. Keihänen, S. Kermiche, A. Kiessling, B. Kubik, M. Kümmel, M. Kunz, H. Kurki-Suonio, Q. Le Boulc'h, A. M. C. Le Brun, D. Le Mignant, S. Ligori, P. B. Lilje, V. Lindholm, I. Lloro, G. Mainetti, D. Maino, E. Maiorano, O. Mansutti, S. Marcin, O. Marggraf, M. Martinelli, N. Martinet, F. Marulli, R. Massey, S. Maurogordato, H. J. McCracken, E. Medinaceli, S. Mei, M. Melchior, Y. Mellier, M. Meneghetti, E. Merlin, G. Meylan, A. Mora, M. Moresco, L. Moscardini, R. Nakajima, C. Neissner, S. -M. Niemi, J. W. Nightingale, C. Padilla, S. Paltani, F. Pasian, K. Pedersen, W. J. Percival, V. Pettorino, S. Pires, G. Polenta, M. Poncet, L. A. Popa, L. Pozzetti, F. Raison, A. Renzi, J. Rhodes, G. Riccio, E. Romelli, M. Roncarelli, R. Saglia, Z. Sakr, A. G. Sánchez, D. Sapone, B. Sartoris, J. A. Schewtschenko, P. Schneider, T. Schrabback, M. Scodeggio, A. Secroun, G. Seidel, M. Seiffert, S. Serrano, P. Simon, C. Sirignano, G. Sirri, L. Stanco, J. Steinwagner, P. Tallada-Crespí, A. N. Taylor, I. Tereno, S. Toft, R. Toledo-Moreo, F. Torradeflot, I. Tutusaus, L. Valenziano, J. Valiviita, T. Vassallo, G. Verdoes Kleijn, A. Veropalumbo, Y. Wang, J. Weller, A. Zacchei, G. Zamorani, F. M. Zerbi, I. A. Zinchenko, E. Zucca, V. Allevato, M. Ballardini, M. Bolzonella, E. Bozzo, C. Burigana, R. Cabanac, A. Cappi, D. Di Ferdinando, J. A. Escartin Vigo, L. Gabarra, J. Martín-Fleitas, S. Matthew, N. Mauri, R. B. Metcalf, A. Pezzotta, M. Pöntinen, C. Porciani, I. Risso, V. Scottez, M. Sereno, M. Tenti, M. Viel, M. Wiesmann, Y. Akrami, I. T. Andika, S. Anselmi, M. Archidiacono, F. Atrio-Barandela, C. Benoist, K. Benson, D. Bertacca, M. Bethermin, L. Bisigello, A. Blanchard, L. Blot, M. L. Brown, S. Bruton, A. Calabro, B. Camacho Quevedo, F. Caro, C. S. Carvalho, T. Castro, Y. Charles, F. Cogato, A. R. Cooray, O. Cucciati, S. Davini, F. De Paolis, G. Desprez, A. Díaz-Sánchez, J. J. Diaz, S. Di Domizio, J. M. Diego, P. -A. Duc, A. Enia, Y. Fang, A. G. Ferrari, P. G. Ferreira, A. Finoguenov, A. Fontana, A. Franco, K. Ganga, J. García-Bellido, T. Gasparetto, V. Gautard, E. Gaztanaga, F. Giacomini, F. Gianotti, G. Gozaliasl, M. Guidi, C. M. Gutierrez, A. Hall, W. G. Hartley, S. Hemmati, C. Hernández-Monteagudo, H. Hildebrandt, J. Hjorth, J. J. E. Kajava, Y. Kang, V. Kansal, D. Karagiannis, K. Kiiveri, C. C. Kirkpatrick, S. Kruk, J. Le Graet, L. Legrand, M. Lembo, F. Lepori, G. Leroy, G. F. Lesci, J. Lesgourgues, L. Leuzzi, T. I. Liaudat, A. Loureiro, J. Macias-Perez, G. Maggio, M. Magliocchetti, E. A. Magnier, F. Mannucci, R. Maoli, C. J. A. P. Martins, L. Maurin, M. Miluzio, P. Monaco, C. Moretti, G. Morgante, C. Murray, K. Naidoo, A. Navarro-Alsina, S. Nesseris, F. Passalacqua, K. Paterson, L. Patrizii, A. Pisani, D. Potter, S. Quai, M. Radovich, S. Sacquegna, M. Sahlén, D. B. Sanders, E. Sarpa, A. Schneider, D. Sciotti, D. Scognamiglio, E. Sellentin, L. C. Smith, K. Tanidis, G. Testera, R. Teyssier, S. Tosi, A. Troja, M. Tucci, C. Valieri, A. Venhola, D. Vergani, G. Verza, P. Vielzeuf, N. A. Walton, J. G. Sorce

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Paper submitted as part of the A&A Special Issue `Euclid Quick Data Release (Q1)', 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06518 2025-03-18 cs.LG q-bio.QM stat.ME 78%

Causal Representation Learning from Multimodal Biomedical Observations

Yuewen Sun, Lingjing Kong, Guangyi Chen, Loka Li, Gongxu Luo, Zijian Li, Yixuan Zhang, Yujia Zheng, Mengyue Yang, Petar Stojanov, Eran Segal, Eric P. Xing, Kun Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09010 2025-03-14 cs.RO 78%

HumanoidPano: Hybrid Spherical Panoramic-LiDAR Cross-Modal Perception for Humanoid Robots

Qiang Zhang, Zhang Zhang, Wei Cui, Jingkai Sun, Jiahang Cao, Yijie Guo, Gang Han, Wen Zhao, Jiaxu Wang, Chenghao Sun, Lingfeng Zhang, Hao Cheng, Yujie Chen, Lin Wang, Jian Tang, Renjing Xu

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05774 2025-03-11 cs.LG cs.DB 78%

GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

Theodor Lundqvist, Ludvig Delvret

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 131 pages, 49 figures, 48 tables

详情

展开后加载摘要…

URL PDF HTML 收藏