arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2509.06641 2025-09-16 cs.AI cs.LG 83%

CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning

Zhou-Peng Shou, Zhi-Qiang You, Fang Wang, Hai-Bo Liu

机构 * NoDesk AI Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :omni-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09427 2025-09-12 cs.CV 83%

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, Huafeng Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Physics and Optoelectronic Engineering, Foshan University(佛山大学物理与光电工程学院) School of Instrumentation Science and Optelectronics Engineering, Beihang University(北航仪器科学与光电工程学院) School of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Information Fusion, 2025, 121: 103146

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20454 2025-09-10 cs.LG cs.CL 83%

CoMMIT: Coordinated Multimodal Instruction Tuning

Xintong Li, Junda Wu, Tong Yu, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣地亚哥分校) Adobe Research(Adobe研究) The University of New South Wales(新南威尔士大学) CSIRO’s Data61(澳大利亚联邦科学与工业研究组织的数据61)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06976 2025-09-10 cs.LG cs.AI 83%

A Knowledge-Guided Cross-Modal Feature Fusion Model for Local Traffic Demand Prediction

Lingyu Zhang, Pengfei Xu, Guobin Wu, Jian Liang, Ruiyang Dong, Yunhai Wang, Xuan Song

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 83%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04938 2025-09-08 cs.MM 83%

An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data

Jianlu Wang, Yanan Wang, Tong Liu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20309 2025-09-08 cs.CV 83%

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Yibo Yang, Si Liu

机构 * Beihang University(北航大学) National University of Singapore(国立新加坡大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09556 2025-09-05 cs.CL 83%

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos

机构 * National Technical University of Athens(希腊国家技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03433 2025-09-04 cs.CV 83%

Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

Kaili sun, Xingyu Miao, Bing Zhai, Haoran Duan, Yang Long

机构 * Department of Computer Science, Durham University, UK.(计算机科学系,杜ham大学,英国) Computer and Information Sciences, Northumbria University, UK.(计算机与信息科学,北umbria大学,英国) Department of Automation, Tsinghua University, China.(自动化系,清华大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00622 2025-09-03 cs.AI cs.IR 83%

BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

Shiqiao Zhou, Holger Schöner, Huanbo Lyu, Edouard Fouché, Shuo Wang

机构 * University of Birmingham(伯明翰大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06989 2025-08-29 cs.CR cs.CV 83%

Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application

Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05849 2025-08-26 cs.CV 83%

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, Cheng-Lin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室) School of Artificial Intelligence, UCAS(人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13460 2025-08-20 cs.CV 83%

Revisiting MLLM Token Technology through the Lens of Classical Visual Coding

Jinming Liu, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Xudong Yang, Zhibo Chen, Huan Wang, Xin Jin

机构 * Eastern Institute of Technology, Ningbo, China(东部技术研究所) Westlake University(西湖大学) USTC(中国科学技术大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13072 2025-08-19 cs.AI 83%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12917 2025-08-19 cs.CV 83%

CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction

Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo, Jie Yang, Yuming Fang, Wei Liu

机构 * School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所,上海交通大学) School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(计算机与人工智能学院,江西财经大学) School of Automation and Intelligent Sensing & Institute of Image Processing and Pattern Recognition & Institute of Medical Robotics, Shanghai Jiao Tong University(自动化与智能感知学院及图像处理与模式识别研究所及医学机器人研究所,上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments The Paper is Accepted by TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07264 2025-08-18 cs.SI cs.AI 83%

FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning

Van Duc Cuong, Ta Dinh Tam, Tran Duc Chinh, Nguyen Thi Hanh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Faculty of Interdisciplinary Digital Technology(跨学科数字技术学院) PHENIKAA University(PHENIKAA大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04058 2025-08-18 cs.CV 83%

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06293 2025-08-15 cs.CV 83%

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness

Qifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu, Bosheng Qin, Wenqiao Zhang, Yunfei Li, Juncheng Li, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09085 2025-08-13 cs.NI cs.AI cs.LG 83%

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

Zihan Fang, Zheng Lin, Senkang Hu, Yihang Tao, Yiqin Deng, Xianhao Chen, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08589 2025-08-13 cs.CV 83%

DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding

Wenwen Yu, Zhibo Yang, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07803 2025-08-12 cs.CV 83%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07681 2025-08-12 cs.LG cs.AI 83%

MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation

Yooseok Lim, ByoungJun Jeon, Seong-A Park, Jisoo Lee, Sae Won Choi, Chang Wook Jeong, Ho-Geol Ryu, Hongyeol Lee, Hyun-Lim Yang

机构 * Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 83%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04571 2025-08-07 cs.IR cs.CL cs.LG 83%

Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation

Claudio Pomo, Matteo Attimonelli, Danilo Danese, Fedelucio Narducci, Tommaso Di Noia

机构 * Sapienza University of Rome(罗马大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted as Full Research Papers at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03102 2025-08-06 cs.CV 83%

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

Tianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02525 2025-08-05 cs.AI 83%

Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model

Qifan Chen, Jin Cui, Cindy Duan, Yushuo Han, Yifei Shi

机构 * King’s College London(伦敦国王学院) Imperial College London(帝国理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01644 2025-08-05 cs.MM cs.AI cs.CV cs.SD eess.AS 83%

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang, Jiaye Yang, Lu Liu, Daibing Yao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in ACM Multimedia 2025. 10 pages, 4 figures

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23444 2025-08-01 cs.MM 83%

Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis

Xiang Li, Xianfu Cheng, Xiaoming Zhang, Zhoujun Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08204 2025-07-30 cs.CV 83%

One-stage Modality Distillation for Incomplete Multimodal Learning

Shicai Wei, Yang Luo, Chunbo Luo

机构 * School of Information and Communication Engineering University of Electronic Science and Technology of China(信息与通信工程学院 电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19839 2025-07-29 cs.LG cs.CV 83%

GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning

Tiantian Peng, Yuyang Liu, Shuo Yang, Qiuhe Hong, YongHong Tian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏