arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2508.02133 2025-08-19 cs.HC 82%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 82%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05476 2025-08-08 eess.IV 82%

MM2CT: MR-to-CT translation for multi-modal image fusion with mamba

Chaohui Gong, Zhiying Wu, Zisheng Huang, Gaofeng Meng, Zhen Lei, Hongbin Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01805 2025-08-05 cs.NI 82%

M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

Yongjie Zeng, Hongyang Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00926 2025-08-05 cs.LG 82%

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

Feng Xu, Hui Wang, Yuting Huang, Danwei Zhang, Zizhu Fan

机构 * Feng Xu 1,2(作者1单位) Hui Wang 1(作者1单位) Yuting Huang 3(作者3单位) Danwei Zhang 4(作者4单位) Zizhu Fan 5(作者5单位)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03248 2025-07-30 cs.CV cs.AI cs.CL 82%

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Yiwu Zhong, Zhuoming Liu, Yin Li, Liwei Wang

机构 * The Chinese University of Hong Kong(香港中文大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19628 2025-07-28 cs.CV cs.CL cs.LG cs.MM 82%

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Qiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye, Zhanpeng Zen, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22334 2025-07-24 cs.CL cs.AI cs.CV cs.LG 82%

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

Lai Wei, Yuting Li, Kaipeng Zheng, Chen Wang, Yue Wang, Linghe Kong, Lichao Sun, Weiran Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23115 2025-07-01 cs.CV cs.AI cs.CL 82%

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Haonan Chen, Hong Liu, Yuping Luo, Liang Wang, Nan Yang, Furu Wei, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Stanford University(斯坦福大学) Microsoft Corporation(微软公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Homepage: https://haon-chen.github.io/MoCa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18023 2025-06-26 cs.CV cs.AI cs.CL 82%

PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding

Kui Huang, Xinrong Chen, Wenyu Lv, Jincheng Liao, Guanzhong Wang, Yi Liu

机构 * Baidu Inc.(百度公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21755 2025-06-24 cs.CV cs.AI cs.CL cs.LG 82%

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering

Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16744 2025-06-23 cs.LG cs.RO eess.SP 82%

IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification

Eion Tyacke, Kunal Gupta, Jay Patel, Rui Li

机构 * Dept. of Electrical and Computer Engineering(电气与计算机工程系) New York University(纽约大学) Tandon School of Engineering(坦顿工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10452 2025-06-13 cs.CV cs.CL cs.LG cs.MM 82%

Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts

Guowei Zhong, Ruohong Huan, Mingzhen Wu, Ronghua Liang, Peng Chen

机构 * College of Computer Science and Technology, Zhejiang University of Technology(浙江工业大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Submitted to TAC. The code is available at https://github.com/gw-zhong/CIDer

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01067 2025-06-12 cs.AI cs.CL cs.CV cs.HC cs.LG 82%

Human-like object concept representations emerge naturally in multimodal large language models

Changde Du, Kaicheng Fu, Bincheng Wen, Yi Sun, Jie Peng, Wei Wei, Ying Gao, Shengpei Wang, Chuncheng Zhang, Jinpeng Li, Shuang Qiu, Le Chang, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Neuroscience, State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology(神经科学研究所) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Chinese Academy of Sciences(中国科学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published on Nature Machine Intelligence

Journal ref Nature Machine Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18956 2025-06-11 cs.CV cs.AI cs.LG cs.MM 82%

How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation

Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息与通信研究院(I2R))

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted at the 2025 International Conference on Machine Learning (ICML)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20188 2025-05-27 cs.LG cs.IR 82%

Research on feature fusion and multimodal patent text based on graph attention network

Zhenzhen Song, Ziwei Liu, Hongji Li

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18536 2025-05-27 cs.CL cs.AI cs.CV 82%

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

Haoyuan Sun, Jiaqi Wu, Bo Xia, Yifu Luo, Yifei Zhao, Kai Qin, Xufei Lv, Tiantian Zhang, Yongzhe Chang, Xueqian Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08334 2025-05-22 cs.CV cs.AI cs.IR cs.MM 82%

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval

Yeong-Joon Ju, Ho-Joong Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23333 2025-04-01 cs.IR cs.AI cs.CL cs.CV 82%

Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics

Jing Zhu, Mingxuan Ju, Yozen Liu, Danai Koutra, Neil Shah, Tong Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21775 2025-03-31 cs.CV cs.AI cs.CL cs.GR cs.LG 82%

StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion

Ziyu Guo, Young Yoon Lee, Joseph Liu, Yizhak Ben-Shabat, Victor Zordan, Mubbasir Kapadia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://stylemotif.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20633 2025-03-27 cs.LG 82%

Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning

Sashuai Zhou, Hai Huang, Yan Xia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments 6 pages,3 figures

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16023 2025-03-21 cs.CR 82%

BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models

Zenghui Yuan, Jiawen Shi, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract)

Comments This paper is accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13709 2025-03-19 cs.LG 82%

Multi-modal Time Series Analysis: A Tutorial and Survey

Yushan Jiang, Kanghui Ning, Zijie Pan, Xuyang Shen, Jingchao Ni, Wenchao Yu, Anderson Schneider, Haifeng Chen, Yuriy Nevmyvaka, Dongjin Song

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13383 2025-03-18 cs.CV cs.AI cs.CL cs.LG 82%

Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning

Mengyao Lyu, Yan Li, Huasong Zhong, Wenhao Yang, Hui Chen, Jungong Han, Guiguang Ding, Zhenheng Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments update comparison with sota and analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10726 2025-03-17 cs.LG eess.IV 82%

Prototype-Guided Cross-Modal Knowledge Enhancement for Adaptive Survival Prediction

Fengchun Liu, Linghan Cai, Zhikang Wang, Zhiyuan Fan, Jin-gang Yu, Hao Chen, Yongbing Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09652 2025-03-14 eess.IV 82%

4D-ACFNet: A 4D Attention Mechanism-Based Prognostic Framework for Colorectal Cancer Liver Metastasis Integrating Multimodal Spatiotemporal Features

Zesheng Li, Wei Yang, Yan Su, Yiran Zhu, Yuhan Tang, Haoran Chen, Chengchang Pan, Honggang Qi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 8 pages,6 figures,2 tables,submitted to the 33rd ACM International Conference on Multimedia(ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06313 2025-03-11 cs.CV cs.AI cs.CL cs.LG cs.RO 82%

Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection

Chandan Kumar Sah, Ankit Kumar Shaw, Xiaoli Lian, Arsalan Shahid Baig, Tuopu Wen, Kun Jiang, Mengmeng Yang, Diange Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00037 2025-03-04 cs.CL cs.AI cs.CV cs.LG 82%

Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs

Wei Zhao, Zhe Li, Yige Li, Jun Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03895 2025-03-04 cs.CV cs.AI cs.CL 82%

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICLR 2025. Code: https://github.com/ictnlp/LLaVA-Mini Model: https://huggingface.co/ICTNLP/llava-mini-llama-3.1-8b

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11712 2025-02-04 cs.IR 82%

Fine-tuning Multimodal Large Language Models for Product Bundling

Xiaohao Liu, Jie Wu, Zhulin Tao, Yunshan Ma, Yinwei Wei, Tat-seng Chua

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted by KDD 2025 (CR)

详情

展开后加载摘要…

URL PDF HTML 收藏