arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2508.15304 2026-01-27 cs.IR 82%

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

MLLMRec: 基于图细化的多模态推荐偏好推理范式

Yuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan, Wanyu Chen, Fei Cai, Honghui Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

AI总结 MLLMRec通过图细化和多模态大语言模型提升多模态推荐的用户偏好推理与物品表示学习准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00670 2026-01-05 cs.HC 82%

Wave2Word: A Multimodal Transformer Framework for Joint EEG-Text Alignment and Multi-Task Representation Learning in Neurocritical Care

Wave2Word: 一种多模态Transformer框架,用于神经重症监护中的联合EEG-文本对齐和多任务表示学习

Argha Kamal Samanta, Deepak Mewada, Monalisa Sarma, Debasis Samanta

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Wave2Word提出一种多模态Transformer框架,通过整合信号域建模与结构化临床语言监督,实现EEG-文本对齐和多任务表示学习,提升神经重症监护中的EEG分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 82%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18686 2025-12-22 cs.LG 82%

Hierarchical Multimodal LLMs with Semantic Space Alignment for Enhanced Time Series Classification

具有语义空间对齐的层次多模态大语言模型用于增强的时间序列分类

Xiaoyu Tao, Tingyue Pan, Mingyue Cheng, Yucong Luo, Qi Liu, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 HiTime通过层次多模态大语言模型和语义空间对齐,提升时间序列分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22410 2025-12-09 stat.AP 82%

Multimodal Fusion and Interpretability in Human Activity Recognition: A Reproducible Framework for Sensor-Based Modeling

多模态融合与可解释性在人体活动识别中的应用:一种可复现的基于传感器建模框架

Yiyao Yang, Yasemin Gulbahar

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)

AI总结 本文提出了一种可复现的多模态融合框架,通过统一预处理和融合策略提升人体活动识别的准确性和可解释性。

Comments 33 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16990 2025-11-24 cs.HC 82%

Senti-iFusion: An Integrity-centered Hierarchical Fusion Framework for Multimodal Sentiment Analysis under Uncertain Modality Missingness

Senti-iFusion: 一种以完整性为中心的多模态情感分析多模态融合框架,用于在不确定模态缺失情况下

Liling Li, Guoyang Xu, Xiongri Shen, Zhifei Xu, Yanbo Zhang, Zhiguo Zhang, Zhenxi Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Senti-iFusion提出了一种以完整性为中心的多模态融合框架,通过分层结构处理模态缺失问题,提升多模态情感分析的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00374 2025-11-18 cs.CV cs.AI cs.MM 82%

MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention

Tianyi Wang, Jianan Fan, Dingxin Zhang, Dongnan Liu, Yong Xia, Heng Huang, Weidong Cai

机构 * The University of Sydney(悉尼大学) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) School of Computer Science and Engineering, Northwestern Polytechnical University(西北工业大学计算机科学与工程学院) University of Maryland(马里兰大学) Ningbo Institute of Northwestern Polytechnical University(西北工业大学宁波学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by IEEE Transactions on Medical Imaging (TMI). Code available at https://github.com/TianyiFranklinWang/MIRROR. Project page: https://tianyifranklinwang.github.io/MIRROR

Journal ref IEEE Trans. Med. Imaging (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10568 2025-11-17 cs.LG cs.AI cs.CL cs.CV 82%

MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Ruixiang Jiang, Lingbo Liu, Changwen Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Research Institute of Multiple Agents and Embodied Intelligence, Pengcheng Laboratory(多智能体与具身智能研究院,鹏城实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to IEEE TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07274 2025-11-11 cs.LG 82%

Multi-modal Dynamic Proxy Learning for Personalized Multiple Clustering

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Ziyue Peng, Zewei Liu, Hewei Wang, Jiayi Zhang, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03371 2025-11-06 cond-mat.mtrl-sci physics.comp-ph 82%

Enhancing composition-based materials property prediction by cross-modal knowledge transfer

Ivan Rubtsov, Ivan Dudakov, Yuri Kuratov, Vadim Korolev

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

Comments 7 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 82%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00987 2025-11-04 cs.LG 82%

Balanced Multimodal Learning via Mutual Information

Rongrong Xie, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究学院(SISSA))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11194 2025-10-27 cs.CE 82%

Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive Alignment

Xiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine, Lawrence P. Petalidis, Michalis Vazirgiannis

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20736 2025-10-24 cs.LG 82%

Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process

Tsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin, Lequan Yu

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by NeruIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20540 2025-10-24 cs.LG 82%

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis

机构 * Center for Wireless Communications, University of Oulu(无线通信中心,奥卢大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 5 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20256 2025-10-24 cs.CV cs.CL cs.LG cs.MM 82%

Calibrating Multimodal Consensus for Emotion Recognition

Guowei Zhong, Junjie Li, Huaiyu Zhu, Ruohong Huan, Yun Pan

机构 * College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学) Zhejiang University Jinhua Research Institute(浙江大学金华研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16824 2025-10-21 cs.LG q-bio.MN 82%

ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning

Yingxu Wang, Kunyu Zhang, Jiaxin Huang, Nan Yin, Siwei Liu, Eran Segal

机构 * MBZUAI(穆扎芬人工智能研究所) University of Zhengzhou(郑州大学) HKUST(香港科技大学) University of Aberdeen(爱丁堡大学) MBZUAI, Weizmann Institute of Science(穆扎芬人工智能研究所、威斯曼科学研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16350 2025-10-21 cs.LG 82%

MGTS-Net: Exploring Graph-Enhanced Multimodal Fusion for Augmented Time Series Forecasting

Shule Hao, Junpeng Bao, Wenli Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12395 2025-10-15 cs.CR 82%

IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion

Ye Tian, Yanqiu Yu, Liangliang Song, Zhiquan Liu, Yanbin Wang, Jianguo Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15595 2025-10-14 cs.RO 82%

Grasping Deformable Objects via Reinforcement Learning with Cross-Modal Attention to Visuo-Tactile Inputs

Yonghyun Lee, Sungeun Hong, Min-gu Kim, Gyeonghwan Kim, Changjoo Nam

机构 * Dept. of Electronic Engineering at Sogang University(ソガン大学电子工程系) Dept. of Immersive Media and Engineering at Sungkyunkwan University(顺天大学沉浸媒体与工程系) College of Medicine, Yonsei University(延世大学医学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08022 2025-10-10 cs.LG cs.AI cs.CL cs.CV 82%

Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining

Chenxi Liu, Tianyi Xiong, Yanshuo Chen, Ruibo Chen, Yihan Wu, Junfeng Guo, Tianyi Zhou, Heng Huang

机构 * University of Maryland, College Park(马里兰大学学院 park)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05283 2025-10-08 cs.AI cs.CL cs.CV 82%

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

Radha Gulhane, Sathish Reddy Indurthi

机构 * Radha Gulhane(独立研究者) Sathish Reddy Indurthi(独立研究者)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09747 2025-10-07 cs.NE 82%

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25278 2025-10-01 cs.LG stat.ML 82%

MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series

Payal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia, Qi Zhu

机构 * Northwestern University(西北大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted to Neurips 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 82%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09114 2025-09-12 cs.IR 82%

Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02943 2025-09-04 cs.IR 82%

Knowledge graph-based personalized multimodal recommendation fusion framework

Yu Fang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20279 2025-08-29 cs.CV cs.AI cs.CL 82%

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Zhuoran Yu, Yong Jae Lee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16147 2025-08-25 cs.IR 82%

Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction

Ao Zhou, Mingsheng Tu, Luping Wang, Tenghao Sun, Zifeng Cheng, Yafeng Yin, Zhiwei Jiang, Qing Gu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13971 2025-08-20 eess.AS cs.CL cs.HC cs.LG cs.MM 82%

Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Andrew Chang, Chenkai Hu, Ji Qi, Zhuojian Wei, Kexin Zhang, Viswadruth Akkaraju, David Poeppel, Dustin Freeman

机构 * New York UniversityUSA(纽约大学) Max Planck SocietyGermany(马克斯·普朗克研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏