arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6847 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6847 篇

2505.06536 2025-05-13 cs.CV cs.AI 88%

TACFN: Transformer-based Adaptive Cross-modal Fusion Network for Multimodal Emotion Recognition

Feng Liu, Ziwang Fu, Yunlong Wang, Qijian Zheng

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) MTlab, Meitu (China) Limited(美图(中国)有限公司) Institute of Acoustics, University of Chinese Academy of Sciences(中国科学院声学研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: text overlap with arXiv:2111.02172

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09707 2025-04-15 cs.AI cs.IT cs.LG cs.MM math.IT 88%

InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing Signals

Tomoyoshi Kimura, Xinlin Li, Osama Hanna, Yatong Chen, Yizhuo Chen, Denizhan Kara, Tianshi Wang, Jinyang Li, Xiaomin Ouyang, Shengzhong Liu, Mani Srivastava, Suhas Diggavi, Tarek Abdelzaher

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18817 2025-03-25 cs.CV cs.AI 88%

Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations

Jeonghyeon Kim, Sangheum Hwang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16629 2025-01-29 cs.CL cs.CV 88%

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs

Jinlan Fu, Shenzhen Huangfu, Hao Fei, Xiaoyu Shen, Bryan Hooi, Xipeng Qiu, See-Kiong Ng

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20821 2024-12-31 eess.AS cs.CL cs.SD 88%

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang, Jiaming Zhou, Yong Qin

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、eess.AS

Comments ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00833 2024-12-03 cs.CV cs.AI 88%

AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

Yan Li, Yifei Xing, Xiangyuan Lan, Xin Li, Haifeng Chen, Dongmei Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16166 2024-10-22 cs.CV cs.CL 88%

Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining

Han Huang, Yuqi Huo, Zijia Zhao, Haoyu Lu, Shu Wu, Bingning Wang, Qiang Liu, Weipeng Chen, Liang Wang

专题命中 多模态训练与对齐 :image-text(title,abstract);MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06610 2024-08-14 cs.CV cs.CL cs.LG 88%

CROME: Cross-Modal Adapters for Efficient Multimodal LLM

Sayna Ebrahimi, Sercan O. Arik, Tejas Nama, Tomas Pfister

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10105 2024-07-16 cs.CV cs.AI 88%

Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification

Tengfei Liu, Yongli Hu, Junbin Gao, Yanfeng Sun, Baocai Yin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00894 2024-01-03 cs.LG cs.CV cs.MM 88%

Balanced Multi-modal Federated Learning via Cross-Modal Infiltration

Yunfeng Fan, Wenchao Xu, Haozhao Wang, Jiaqi Zhu, Song Guo

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(title);multimodal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 5 figures 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14302 2023-10-17 cs.IR cs.AI cs.HC cs.LG cs.MM 88%

VIP5: Towards Multimodal Foundation Models for Recommendation

Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu, Yongfeng Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07646 2023-06-14 cs.CV cs.MM 88%

Enhanced Multimodal Representation Learning with Cross-modal KD

Mengxi Chen, Linyu Xing, Yu Wang, Ya Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02646 2022-11-07 cs.LG cs.AI cs.MM 88%

Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions

Gaurav Verma, Vishwa Vinay, Ryan A. Rossi, Srijan Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted at the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP); Full Paper (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.10421 2020-10-26 cs.CL cs.IR cs.MM 88%

Multimodal Analytics for Real-world News using Measures of Cross-modal Entity Consistency

Eric Müller-Budack, Jonas Theiner, Sebastian Diering, Maximilian Idahl, Ralph Ewerth

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted for publication in: International Conference on Multimedia Retrieval (ICMR), Dublin, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25422 2026-07-29 cs.AI 新提交 88%

Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

显著知识路径:用于高效知识密集型多模态问答的稀疏跨模态路由

Noor Islam S. Mohammad, Uluğ Bayazıt

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 研究知识密集型多模态问答,提出SKIP架构,通过问题引导视觉令牌修剪等方法,沿稀疏路径计算路由,结合自适应预算控制器,在五个基准测试中,以更少计算量和更低延迟达到或超越密集基线准确性。

Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026) Workshop on Efficient Multimodal Question Answering (EMM-QA), Seoul, South Korea. Copyright 2026 by the author(s). (Archival)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18780 2026-06-18 cs.CV cs.CL cs.MM 新提交 88%

SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction

SAMA:面向统一低资源多模态信息抽取的语义锚定对齐增强

Quanjiang Guo, Chong Mu, Jiazhou Pan, Ming Jia, Ling Tian, Hui Gao, Zhao Kang

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 提出语义锚定对齐增强框架SAMA,通过构建结构化语义锚引导多专家多模态大模型生成高保真文本,并利用锚保留扩散机制合成图像,结合双约束过滤模块,在低资源多模态信息抽取任务中显著提升性能。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03450 2026-08-05 cs.MM cs.AI cs.CL cs.CV 新提交 88%

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

平衡效率与效能:面向多模态大语言模型(MLLM)的无训练注意力引导式显式与隐式思维切换

Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang

专题命中 多模态训练与对齐 :MLLM(title_cn,summary_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对MLLM推理中感知与逻辑混淆的问题,提出无训练的注意力引导式切换框架,通过视觉-文本注意力比率自适应切换显式与隐式推理,实现性能与效率的双重提升。

Comments Accepted by ACM MM 2026. 10 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21952 2026-07-07 cs.LG cs.AI cs.AR cs.NE cs.RO 88%

Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models

聚焦会议:加速多模态基础模型的硬件和软件技术

Muhammad Shafique, Abdul Basit, Muhammad Abdullah Hanif, Alberto Marchisio, Rachmad Vidya Wicaksana Putra, Minghao Shao

机构 * eBRAIN Lab, New York University (NYU) Abu Dhabi(eBRAIN实验室,纽约大学(NYU)阿布扎比)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

AI总结 本文提出了一种多层方法,通过硬件和软件协同设计和优化流程,提升多模态基础模型的效率。方法包括混合精度量化、结构剪枝、推测解码等技术,结合硬件加速器优化计算和内存需求,验证了在医疗和代码生成任务中的有效性。

Comments Accepted at the Design, Automation and Test in Europe Conference (DATE), April 20-22, 2026 in Verona, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08774 2026-06-30 cs.IR cs.AI 88%

Multimodal Representation Alignment for Cross-modal Information Retrieval

多模态表示对齐用于跨模态信息检索

Fan Xu, Luis A. Leiva

机构 * University of Luxembourg(卢森堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文研究多模态表示对齐对跨模态检索的影响,通过对比不同度量标准和嵌入空间,发现余弦相似度在对齐任务中表现最佳,同时 Wasserstein 距离提供了分布差异的补充视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26894 2026-06-26 cs.CV 新提交 88%

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

多模态3D MRI中的局部、全局和跨模态上下文建模

Minh Duc Do, Tillmann Rheude, Noel Kronenberg, Roland Eils, Benjamin Wild

机构 * Berlin Institute of Health at Charité - Universitätsmedizin Berlin(柏林健康研究所,柏林夏里特医学院) Health Data Science Unit, Heidelberg University Hospital and BioQuant(海德堡大学医院与BioQuant健康数据科学部) Intelligent Medicine Institute, Fudan University(复旦大学智能医学研究所) Department of Mathematics and Computer Science, Freie Universität Berlin(柏林自由大学数学与计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 提出MICViT,一种3D视觉Transformer,通过四种注意力机制显式建模模态内和跨模态的局部与全局交互,在多数据集脑龄预测任务中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04719 2026-06-04 cs.CL 88%

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

基于查询的跨模态投影器增强Mamba多模态大语言模型

SooHwan Eom, Jay Shim, Gwanhyeong Koo, Haebin Na, Mark A. Hasegawa-Johnson, Sungwoong Kim, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology / Korea, Republic of(韩国科学技术院) University of Illinois in Urbana-Champaign / United States of America(伊利诺伊大学厄巴纳-香槟分校) Korea University / Korea, Republic of(韩国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL

AI总结 提出基于查询的跨模态投影器,通过交叉注意力压缩视觉令牌,消除手动设计2D扫描顺序的需求,提升Mamba多模态LLM的性能和吞吐量。

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16657 2026-04-21 cs.LG cs.AI 88%

Cross-Modal Bayesian Low-Rank Adaptation for Uncertainty-Aware Multimodal Learning

跨模态贝叶斯低秩适应用于不确定性感知的多模态学习

Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin

机构 * Dalhousie University(达尔豪斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出CALIBER框架,通过跨模态注意力机制实现多模态不确定性感知的参数高效微调,实验表明其在文本和音频模型上均优于传统基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11892 2026-04-14 cs.CV 88%

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

DecAlign:解耦多模态表示学习的分层跨模态对齐

Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯A&M大学) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 DecAlign通过分层跨模态对齐框架解耦多模态表示,结合原型引导最优传输策略和多模态Transformer提升语义一致性,实验显示在四个基准上优于现有方法。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08579 2026-04-13 cs.LG cs.AI 88%

On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment

关于跨模态表示的谱几何:一种用于多模态对齐的功能图诊断

Krisanu Sarkar

机构 * Indian Institute of Technology Bombay(印度理工学院孟买分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本文研究了基于功能图框架的跨模态对齐,发现尽管功能图在不同监督预算下表现不如Procrustes对齐,但揭示了多模态表示的结构特性,提出了三种诊断量以表征跨模态表示兼容性。

Comments Under review at ACMMM Brave New Ideas Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01536 2026-03-03 cs.IR cs.MM 88%

CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation

CLEAR: 多模态去冗余中的跨模态空域投影

Hao Zhan, Yihui Wang, Yonghui Yang, Danyang Yue, Yu Wang, Pengyang Shao, Fei Shen, Fei Liu, Le Wu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.MM

AI总结 CLEAR通过显式减少跨模态冗余提升多模态推荐性能,采用子空间投影方法增强表示学习动态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00695 2026-03-03 cs.CV 88%

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

STMI: 基于跨模态超图交互的多模态目标重识别分割引导令牌调节

Xingguo Xu, Zhanyu Liu, Weixiang Zhou, Yuansheng Gao, Junjie Cao, Yuhao Wang, Jixiang Luo, Dell Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 STMI通过分割引导特征调节、语义令牌重新分配和跨模态超图交互,提升多模态目标重识别的性能与鲁棒性。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22405 2026-02-27 cs.LG cs.CV 88%

MolFM-Lite: Multi-Modal Molecular Property Prediction with Conformer Ensemble Attention and Cross-Modal Fusion

MolFM-Lite:基于构象集合注意力和跨模态融合的多模态分子性质预测

Syed Omer Shah, Mohammed Maqsood Ahmed, Danish Mohiuddin Mohammed, Shahnawaz Alam, Mohd Vahaj ur Rahman

机构 * Department of Computer Science and Engineering, University at Buffalo(布法罗大学计算机科学与工程系) Khoury College of Computer Sciences, Northeastern University(东北大学科赫里学院) Department of Computer Science, Muffakham Jah College of Engineering and Technology(穆法卡姆·贾赫工程与技术学院计算机科学系) Department of Computer Science and Artificial Intelligence, Muffakham Jah College of Engineering and Technology(穆法卡姆·贾赫工程与技术学院计算机科学与人工智能系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 MolFM-Lite通过跨模态融合和构象集合注意力机制,提升多模态分子性质预测的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21648 2026-02-10 cs.CV cs.CY cs.HC 88%

CAF-Mamba: Mamba-Based Cross-Modal Adaptive Attention Fusion for Multimodal Depression Detection

CAF-Mamba:基于Mamba的跨模态自适应注意力融合用于多模态抑郁症检测

Bowen Zhou, Marc-André Fiedler, Ayoub Al-Hamadi

机构 * Neuro-Information Technology Group (NIT) IIKT, Otto von Guericke University Magdeburg Magdeburg, Germany(奥托·冯·格里克大学马格德堡分校神经信息技术小组(NIT)IIKT)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

AI总结 CAF-Mamba通过基于Mamba的跨模态自适应注意力融合框架,提升多模态抑郁症检测的性能。

Comments The paper contains a total of 5 pages and 3 figures. This paper has been accepted for publication in the proceedings of 2026 IEEE ICASSP Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23936 2026-01-01 cs.CV 88%

MGML: A Plug-and-Play Meta-Guided Multi-Modal Learning Framework for Incomplete Multimodal Brain Tumor Segmentation

MGML: 一种插件式元引导多模态学习框架用于不完整多模态脑肿瘤分割

Yulong Zou, Bo Liu, Cun-Jing Zheng, Yuan-ming Geng, Siyue Li, Qiankun Zuo, Shuihua Wang, Yudong Zhang, Jin Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(title,abstract);分类 cs.CV

AI总结 MGML框架通过元引导和一致性正则化模块提升不完整多模态脑肿瘤分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15741 2025-11-21 cs.AI cs.HC cs.LG 88%

Uncertainty-Resilient Multimodal Learning via Consistency-Guided Cross-Modal Transfer

通过一致性引导的跨模态转移实现不确定性鲁棒的多模态学习

Hyo-Jeong Jang

机构 * Brain and Cognitive Engineering(脑科学与认知工程)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 本论文提出通过一致性引导的跨模态转移实现不确定性鲁棒的多模态学习,旨在提高模型的稳定性和鲁棒性。

Comments Master's thesis, Korea University, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏