arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2512.19934 2025-12-24 cs.CV cs.AI cs.LG 86%

Vehicle-centric Perception via Multimodal Structured Pre-training

基于多模态结构预训练的车辆感知

Wentao Wu, Xiao Wang, Chenglong Li, Jin Tang, Bin Luo

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province(安徽省信息材料与智能感知实验室) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室) the School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VehicleMAE-V2,通过多模态结构先验知识提升车辆感知的预训练能力,采用SMM、CRM和SRM模块增强模型对车辆结构和语义的理解。

Comments Journal extension of VehicleMAE (AAAI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13070 2025-09-17 cs.CV cs.AI 86%

TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation

Qianqi Lu, Yuxiang Xie, Jing Zhang, Shiwei Zou, Yan Chen, Xidao Luan

专题命中 多模态训练与对齐 :image-text(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01343 2025-08-08 cs.CV cs.AI cs.LG 86%

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

Bingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * University of Science and Technology of China(科学技术大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Northwestern Polytechnical University(西北工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03734 2025-08-07 eess.IV cs.AI cs.CV 86%

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du, Shuo Yang, Meidan Ding, Qihao Xu, Chengliang Liu, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China(深圳大学计算机科学与软件工程学院) Shenzhen Key Laboratory of Visual Object Detection and Recognition, Harbin Institute of Technology, Shenzhen, 518055, China(视觉对象检测与识别深圳重点实验室) Laboratory for Artificial Intelligence in Design, Hong Kong(人工智能设计实验室) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09999 2025-06-13 cs.LG cs.MM cs.SD eess.AS 86%

Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion

Yukun Chen, Zihuan Qiu, Fanman Meng, Hongliang Li, Linfeng Xu, Qingbo Wu

机构 * University of Electronic Science and Technology of China(电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09126 2024-12-13 cs.MM cs.AI cs.LG 86%

Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning

Meng Shen, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu, Simon See

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 11 pages, ACMMM Asia 2024, Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17531 2024-10-29 cs.CV cs.AI 86%

SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

Ming Dai, Lingfeng Yang, Yihao Xu, Zhenhua Feng, Wankou Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 24pages, 18figures, NeurIPS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05938 2024-10-10 cs.CV cs.AI 86%

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

Yifei Xing, Xiangyuan Lan, Ruiping Wang, Dongmei Jiang, Wenjun Huang, Qingfang Zheng, Yaowei Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02906 2024-06-18 cs.CR cs.CL cs.CV 86%

MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Renjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, Tong Zhang

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02030 2024-06-06 cs.CL cs.AI 86%

Multimodal Reasoning with Multimodal Knowledge Graph

Junlin Lee, Yequan Wang, Jing Li, Min Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18959 2024-05-30 cs.CV cs.MM 86%

Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval

Rui Yang, Shuang Wang, Yingping Han, Yuanheng Li, Dong Zhao, Dou Quan, Yanhe Guo, Licheng Jiao

专题命中 多模态训练与对齐 :image-text(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11338 2024-05-24 cs.CV cs.AI 86%

EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging

Danli Shi, Weiyi Zhang, Xiaolan Chen, Yexin Liu, Jiancheng Yang, Siyu Huang, Yih Chung Tham, Yingfeng Zheng, Mingguang He

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11311 2024-03-26 cs.CL cs.MM 86%

Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding

Zichen Wu, Hsiu-Yuan Huang, Fanyi Qu, Yunfang Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.MM

Comments LREC-COLING 2024, Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13592 2023-08-22 cs.MM cs.LG cs.SD eess.AS 86%

TACOformer:Token-channel compounded Cross Attention for Multimodal Emotion Recognition

Xinda Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted by IJCAI 2023- AI4TS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07664 2022-08-23 cs.MM cs.CV 86%

M2HF: Multi-level Multi-modal Hybrid Fusion for Text-Video Retrieval

Shuo Liu, Weize Quan, Ming Zhou, Sihong Chen, Jian Kang, Zhe Zhao, Chen Chen, Dong-Ming Yan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM

Comments 1 1pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00693 2021-09-03 cs.CV cs.AI 86%

AnANet: Modeling Association and Alignment for Cross-modal Correlation Classification

Nan Xu, Junyan Wang, Yuan Tian, Ruike Zhang, Wenji Mao

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.13198 2021-04-23 cs.CL cs.CV 86%

InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, Hongxia Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24327 2026-04-02 cs.CV 86%

Le MuMo JEPA: Multi-Modal Self-Supervised Representation Learning with Learnable Fusion Tokens

Le MuMo JEPA:多模态自监督表示学习中的可学习融合标记

Ciem Cornelissen, Sam Leroux, Pieter Simoens

机构 * IDLab, Department of Information Technology, Ghent University - imec(根特大学-imec信息技术系IDLab)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract,comments);multimodal(abstract);分类 cs.CV

AI总结 本文提出Le MuMo JEPA框架,通过学习融合标记实现多模态统一表示学习,实验显示其在性能与效率之间取得最佳平衡,尤其在CenterNet检测和密集深度估计中表现优异。

Comments 14 pages, 4 figures, supplementary material. Accepted at the CVPR 2026 Workshop on Unified Robotic Vision with Cross-Modal Sensing and Alignment (URVIS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01850 2026-06-08 cs.CV cs.AI cs.LG cs.MM 版本更新 86%

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs

MoDA: 面向指令型多模态大语言模型的细粒度视觉定位的调制适配器

Wayner Barrios, Andrés Villa, Juan León Alcázar, SouYoung Jin, Bernard Ghanem

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 提出MoDA调制适配器,通过指令引导的通道级乘法调制增强细粒度视觉定位,在12个基准上对三种MLLM架构取得一致提升,计算开销极小。

Comments Accepted at ICML 2026. Code is available at https://github.com/waybarrios/MoDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19030 2026-01-22 stat.ME math.ST stat.TH 86%

Multimodal data integration and cross-modal querying via orchestrated approximate message passing

多模态数据整合与跨模态查询 via 协调近似消息传递

Sagnik Nandy, Zongming Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title)

AI总结 本文提出了一种协调近似消息传递算法,用于多模态数据整合与跨模态查询,通过统计最优信号恢复和渐近有效的预测集构建,提升单细胞数据的分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24693 2025-09-30 q-bio.NC 86%

Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens

Zijian Dong, Ruilin Li, Joanna Su Xian Chong, Niousha Dehestani, Yinghui Teng, Yi Lin, Zhizhou Li, Yichi Zhang, Yapei Xie, Leon Qi Rong Ooi, B. T. Thomas Yeo, Juan Helen Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title)

Comments NeurIPS 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08199 2026-08-14 cs.CV 版本更新 85%

Generalizable Operating Room Expert with Multimodal Enhancement

具备多模态增强能力的可泛化手术室专家

Peiqi He, Zhenhao Zhang, Yixiang Zhang, Jiaxin Liu, Xiongjun Zhao, Shaoliang Peng

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 针对手术室空间建模的局限,提出仅用RGB图像推理的多模态大语言模型OR-Expert,通过内部推导空间线索实现三维推理,在手术室基准上达SOTA且泛化性良好。

Comments Accepted by ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05967 2026-08-07 cs.MM 新提交 85%

M$^3$Prune: Hierarchical Collaborative Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

M³Prune:面向高效多模态多智能体检索增强生成的分层协同剪枝

Taolin Zhang, Weizi shao, Zijie Zhou, Chen Chen, Daiyang Yu, Tingyuan Hu, Chengyu Wang, Xiaofeng He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.MM

AI总结 M³Prune是优化多模态多智能体检索增强生成的分层协同剪枝框架,通过剪枝冗余通信边提升性能与token效率,实验中优于单智能体及多智能体基准系统。

Comments Accepted by ACM MM2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04013 2026-08-06 cs.LG cs.AI 新提交 85%

C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning

C²MOE:面向不完整多模态情感学习的一致性与互补性引导的混合专家模型

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

机构 * Central South University of Forestry and Technology(中南林业科技大学) State University of New York, New Paltz(纽约州立大学新帕尔茨分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 针对多模态情感识别中模态缺失导致性能下降的问题,本文提出C²MOE框架,通过分解多模态知识为一致性与互补性分量、引入双分支预测机制及可学习重加权模块,在多个基准上实现了优于现有方法的性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03812 2026-08-05 cs.CV 新提交 85%

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

OmniPack:面向高效全模态大语言模型的统一令牌压缩方法

Wanshun Su, Yang Shi, Feihu Liu, Ziwen Yu, Yan Min, Zhuoran Zhang, Qixun Wang, Haotian Wang, Shixuan Liu, Yuanxing Zhang, Peng Wu, Chengfu Huo, Liang Ding

专题命中 多模态训练与对齐 :omni-modal(title,abstract);multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文针对全模态大语言模型的高计算开销问题,提出无需训练的OmniPack框架,通过LLM前结构压缩与LLM内语义优化的协作,在5个基准上实现最优性能-效率权衡,在Qwen2.5-Omni-7B上大幅降低计算量且保持高性能。

Comments 16 pages, 5 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28006 2026-07-31 cs.AI 新提交 85%

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

MMLDSum-LLM:结合视觉对齐与关键词感知的多模态长文档摘要方法

Xianpeng Zhang, Jiahua Yang, Dongyu Chen, Lei zhang, Jian Ma, Xu guohuan, Haonan Lu, Tianhuang Su, Chuangchuang Wang, Kai Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

AI总结 针对多模态长文档摘要的关键信息遗漏与跨模态幻觉问题,提出结合视觉对齐与关键词感知的两阶段训练框架MMLDSum-LLM,在自研基准MMLDSum-Bench上验证了其性能优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26885 2026-07-30 cs.CV 新提交 85%

SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

SCALPEL:基于大语言模型驱动的编码器学习的医学视觉-语言语义跨模态对齐框架

Yunzhan Fu, Enyu Bao, Xiangyu Shen, Yihao Wu, Chunbo Jiang, Fangli Guan, Liqi Yan

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 针对现有医学视觉-语言预训练的三大瓶颈,本文提出SCALPEL框架,通过临床报告对比微调、非对称对齐策略及解剖-否定感知目标,在三大医学基准上实现了多项任务的最优性能。

Comments 14 pages, 3 figures, accepted by PRCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13698 2026-07-15 cs.CV 版本更新 85%

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

注意力忽视视觉风险:多模态安全对齐的风险适应性转向

Jonghyun Park, Minhyuk Seo, Chaewon Yeo, Jonghyun Choi

机构 * Seoul National University(首尔大学) KU Leuven(鲁汶大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MoRAS,通过简洁的视觉上下文增强关键安全区域的视觉注意力,以实现多模态安全对齐,减少推理开销并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11030 2026-07-14 cs.IR cs.LG cs.MM 新提交 85%

MMRM: A Multiplex Multimodal Representation Model for Product Ranking in E-commerce Search

MMRM:一种用于电子商务搜索中产品排名的多重多模态表示模型

Zhen-Lin Chen, Maosen Sheng, Peng Lin, Jianmin Chen, Zhuojian Xiao, Dongyue Wang, Xiwei Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.MM

AI总结 针对电子商务搜索排名中多模态信息利用的局限,提出MMRM框架,通过共享主干从多信号学习生成多重商品表示,引入多重用户表示策略,并经实验验证其有效性,已成功应用于京东搜索引擎提升性能。

Comments Accepted by SIGIR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23041 2026-07-03 cs.CV 新提交 85%

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

SPAR: 语义-像素自对齐与自适应路由的统一多模态模型

Hongxiang Li, Hongxu Chen, Chenyang Zhu, Xiaoshuang Huang, Jiayin Cai, Xiaolong Jiang, Yao Hu, Long Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学) Xiaohongshu Inc.(小红书公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出SPAR框架,通过非对称双流统一分词器协调语义感知与像素重建,利用自对齐生成范式消除外部依赖,并引入动态令牌路由实现灵活多模态交互,在统一架构中达到生成、重建与视觉理解的最优性能。

Comments ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏