arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2511.01320 2025-11-04 cs.AI 83%

OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance

Ziqi Wang, Hailiang Zhao, Yuhao Yang, Daojiang Hu, Cheng Bao, Mingyi Liu, Kai Di, Schahram Dustdar, Zhongjie Wang, Shuiguang Deng

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Hangzhou School of Automation, Zhejiang Normal University(浙江师范大学杭州自动化学院) Distributed Systems Group at the TU Wien and with ICREA at the UPF, Barcelona(维也纳大学分布式系统组和巴塞罗那大学ICREA)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27091 2025-11-03 cs.LG cs.AI quant-ph 83%

QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis

Yiwei Chen, Kehuan Yan, Yu Pan, Daoyi Dong

机构 * School of Engineering, Yunnan University(云南大学工程学院) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Institute of Cyber-Systems and Control, College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院智能系统与控制研究所) Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23640 2025-10-29 cs.LG cs.AI 83%

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18672 2025-10-29 cs.CV 83%

CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion

Juncen Guo, Siao Liu, Xiaoguang Zhu, Lianlong Sun, Liangyu Teng, Jingyi Wu, Di Li, Linxiao Gong, Weiwei Jiang, Wei Zhou, Liang Song

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(智能机器人与先进制造学院,复旦大学) School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) DataLab: Data Science and Informatics, University of California, Davis(数据实验室:数据科学与信息学,加州大学戴维斯分校) University of Rochester(罗切斯特大学) Ningbo University(宁波大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Beijing University of Posts and Telecommunications(北京邮电大学) Academy for Computer Science and Informatics, Cardiff University(计算机科学与信息学学院,卡迪夫大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23273 2025-10-28 cs.LG cs.AI q-bio.QM 83%

A Novel Framework for Multi-Modal Protein Representation Learning

Runjie Zheng, Zhen Wang, Anjie Qiao, Jiancong Xie, Jiahua Rao, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University (SYSU)(计算机科学与工程学院,中山大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 35 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23151 2025-10-28 cs.CV cs.LG 83%

AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes

Sixian Liu, Chen Xu, Qiang Wang, Donghai Shi, Yiwen Li

机构 * Yaowu Technology Co., Ltd, Shenzhen, China(深圳优华科技有限公司,深圳,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21829 2025-10-28 cs.CV 83%

A Flow Model with Low-Rank Transformers for Incomplete Multimodal Survival Analysis

Yi Yin, Yuntao Shou, Zao Dai, Yun Peng, Tao Meng, Wei Ai, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业与技术大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01890 2025-10-27 cs.LG cs.AI 83%

CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection

David Ortiz-Perez, Manuel Benavent-Lledo, Javier Rodriguez-Juan, Jose Garcia-Rodriguez, David Tomás

机构 * Department of Computer Science and Technology, University of Alicante, Alicante, Spain(计算机科学与技术系,阿利坎特大学,阿利坎特,西班牙) Valencian Graduate School and Research Network of Artificial Intelligence, Valencia, Spain(瓦伦西亚人工智能研究生学校与研究网络,瓦伦西亚,西班牙) Department of Software and Computing Systems, University of Alicante, Alicante, Spain(软件与计算系统系,阿利坎特大学,阿利坎特,西班牙)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Journal ref Knowledge-Based Systems, Vol. 329, 2025, Article 114264

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14349 2025-10-24 cs.CV 83%

Vision-Centric Activation and Coordination for Multimodal Large Language Models

Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang, Ziqiang Li, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(人工智能MoE实验室,上海交通大学) Ant Group(蚂蚁集团) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究所,东部技术研究所,宁波)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19336 2025-10-23 cs.CV 83%

DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents

Kai Shi, Jun Yang, Ni Yang, Binqiang Pan, Qingsong Xie, Chao Zhang, Zhenyu Yang, Tianhuang Su, Haonan Lu

机构 * OPPO AI Center(OPPO人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09385 2025-10-22 cs.CV 83%

ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model

Jialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin, Dongyue Wu, Nong Sang, Liang Pan, Changxin Gao

机构 * National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱信息智能处理技术国家实验室,人工智能与自动化学院,华中科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02477 2025-10-16 cs.RO cs.CV 83%

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng, Lue Fan, Dong An, Changwei Wang, Li Guo, Weiliang Meng, Xiaopeng Zhang, Rongtao Xu, Shibiao Xu

机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts Telecommunications, China [1ex] dKey Laboratory of Computing Power Network Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet Service Computing, Shandong Fundamental Research Center for Computer Science, China

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026

Journal ref Information Fusion, 126 (Part B), February 2026, 103652

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11175 2025-10-14 cs.CV 83%

Reliable Cross-modal Alignment via Prototype Iterative Construction

Xiang Ma, Litian Xu, Lexin Fang, Caiming Zhang, Lizhen Cui

机构 * Shandong University(山东大学) The University of Exeter(埃克塞特大学) The Joint SDU-NTU Centre for Artificial Intelligence Research(SDU-NTU联合人工智能研究中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10679 2025-10-14 cs.CV 83%

MSM-Seg: A Modality-and-Slice Memory Framework with Category-Agnostic Prompting for Multi-Modal Brain Tumor Segmentation

Yuxiang Luo, Qing Xu, Hai Huang, Yuqi Ouyang, Zhen Chen, Wenting Duan

机构 * Graduate School of Information, Production and Systems, Waseda University, Japan(信息、生产与系统研究生院,早稻田大学,日本) School of Computer Science, University of Lincoln, UK(林肯大学计算机科学学院,英国) University of Nottingham, UK(诺丁汉大学,英国) University of Nottingham Ningbo China, China(宁波诺丁汉大学,中国) College of Electrical Engineering and Information, Northeast Agricultural University, Harbin, China(电气工程与信息学院,东北农业大学,中国) College of Computer Science, Sichuan University, Chengdu, China(计算机科学学院,四川大学,中国) Yale University, New Haven, CT 06510, USA(耶鲁大学,美国) School of Engineering and Physical Science, University of Lincoln, Lincoln LN6 7TS, UK(工程与物理科学学院,林肯大学,英国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17040 2025-10-14 cs.CV 83%

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

机构 * Peking University(北京大学) Northeastern University(东北大学) Sydney Smart Technology College(悉尼智能技术学院) School of Computer Science, Peking University(北京大学计算机学院) The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IJCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07326 2025-10-10 cs.MM cs.SD 83%

Audio-Visual Separation with Hierarchical Fusion and Representation Alignment

Han Hu, Dongheng Lin, Qiming Huang, Yuqi Hou, Hyung Jin Chang, Jianbo Jiao

专题命中 多模态训练与对齐 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06113 2025-10-08 cs.CV 83%

Multimodal Feature Prototype Learning for Interpretable and Discriminative Cancer Survival Prediction

Shuo Jiang, Zhuwen Chen, Liaoman Xu, Yanming Zhu, Changmiao Wang, Jiong Zhang, Feiwei Qin, Yifei Chen, Zhu Zhu

机构 * Hangzhou Dianzi University(杭州电子科技大学) School of Information and Communication Technology, Griffith University(信息与通信技术学院,格里菲斯大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Laboratory of Advanced Theranostic Materials and Technology, Ningbo Institute of Materials Technology and Engineering, Chinese Academy of Sciences(先进诊疗材料与技术实验室,宁波材料技术与工程研究所,中国科学院) School of Biomedical Engineering, Tsinghua University(生物医学工程学院,清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02086 2025-10-08 cs.LG cs.CV stat.ML 83%

Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations

Minoh Jeong, Zae Myung Kim, Min Namgung, Dongyeop Kang, Yao-Yi Chiang, Alfred Hero

机构 * University of Michigan(密歇根大学) University of Minnesota(明尼苏达大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01954 2025-10-03 cs.CV 83%

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu, Jingyi Liao, Junyi Pan, Yuan Liu, Xiaofen Xing, Chong Sun, Chen Li, Nancy F. Chen, Shuicheng Yan, Xulei Yang, Xun Xu

机构 * South China University of Technology(华南理工大学) Institute for Infocomm Research (I 2 R), A*STAR(信息通信研究所(I 2 R),A*STAR) WeChat Vision, Tencent Inc.(微信视觉,腾讯公司) Foshan University(佛山大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 24 pages, 12 figures and 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01677 2025-10-03 cs.LG cs.CV 83%

Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis

Han Wu, Yanming Sun, Yunhe Yang, Derek F. Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01428 2025-10-03 q-bio.QM cs.AI 83%

BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan, Ben Shapira

机构 * IBM T.J. Watson Research Center(IBM T.J. Watson研究所以) IBM Research(IBM研究所以)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00320 2025-10-03 cs.CV 83%

TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models

Hao Zhang, Mengsi Lyu, Chenrui He, Yulong Ao, Yonghua Lin

机构 * Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24505 2025-09-30 cs.CV 83%

Robust Multimodal Semantic Segmentation with Balanced Modality Contributions

Jiaqi Tan, Xu Zheng, Fangyu Li, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) HKUST(GZ)(香港科技大学(广州)) INSAIT

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20022 2025-09-25 cs.CV 83%

PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival Prediction

Manahil Raza, Ayesha Azam, Talha Qaiser, Nasir Rajpoot

机构 * University of Warwick, UK(沃里克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19628 2025-09-25 cs.CE cs.CL q-fin.CP 83%

Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series

Ross Koval, Nicholas Andrews, Xifeng Yan

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12623 2025-09-24 cs.SD cs.AI cs.CL cs.MM eess.AS 83%

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji

机构 * Sony Group Corporation(索尼集团公司) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 83%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16149 2025-09-22 cs.CV 83%

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu, Jiahui Gao, Jipeng Zhang, Xiaofang Zhou

机构 * HKUST(香港科技大学) HKU(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16017 2025-09-22 cs.CV 83%

DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching

Meng Yang, Fan Fan, Zizhuo Li, Songchu Deng, Yong Ma, Jiayi Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14383 2025-09-19 cs.RO cs.CV 83%

RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings

Yuhong Lu

机构 * Samueli School of Engineering, Electrical and Computer Engineering, UCLA(UCLA电气与计算机工程学院萨姆利学校)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments This paper is submitted to IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏