arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2012.11211 2020-12-22 eess.IV cs.CV 83%

A Multi-View Dynamic Fusion Framework: How to Improve the Multimodal Brain Tumor Segmentation from Multi-Views?

Yi Ding, Wei Zheng, Guozheng Wu, Ji Geng, Mingsheng Cao, Zhiguang Qin

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.03908 2020-11-10 eess.IV cs.CV cs.LG 83%

Cross-Modal Self-Attention Distillation for Prostate Cancer Segmentation

Guokai Zhang, Xiaoang Shen, Ye Luo, Jihao Luo, Zeju Wang, Weigang Wang, Binghui Zhao, Jianwei Lu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05396 2020-10-27 cs.CV cs.LG 83%

Multimodal Self-Supervised Learning for Medical Image Analysis

Aiham Taleb, Christoph Lippert, Tassilo Klein, Moin Nabi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2019 Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01523 2020-09-04 cs.CV 83%

A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports

Yikuan Li, Hanyin Wang, Yuan Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 10 pages, 3 figures, submitted to BIBM2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.00135 2019-09-17 cs.CV 83%

RFBNet: Deep Multimodal Networks with Residual Fusion Blocks for RGB-D Semantic Segmentation

Liuyuan Deng, Ming Yang, Tianyi Li, Yuesheng He, Chunxiang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.05649 2019-08-16 eess.IV cs.CV 83%

A Multimodal Vision Sensor for Autonomous Driving

Dongming Sun, Xiao Huang, Kailun Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.13072 2019-05-01 cs.CV 83%

Cross-Modal Message Passing for Two-stream Fusion

Dong Wang, Yuan Yuan, Qi Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 2018 IEEE International Conference on Acoustics, Speech and Signal Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.06496 2019-03-18 cs.LG cs.CV cs.NE 83%

MFAS: Multimodal Fusion Architecture Search

Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, Frédéric Jurie

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR 2019, Jun 2019, Long Beach, United States http://cvpr2019.thecvf.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.00049 2018-06-19 cs.CV cs.LG 83%

Medical Image Segmentation Based on Multi-Modal Convolutional Neural Network: Study on Image Fusion Schemes

Zhe Guo, Xiang Li, Heng Huang, Ning Guo, Quanzheng Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Zhe Guo and Xiang Li contribute equally to this work

Journal ref 2018 IEEE 15th International Symposium on Biomedical Imaging (ISBI 2018), Washington, DC, 2018, pp. 903-907

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.04503 2017-07-25 cs.CL cs.AI cs.CV cs.MM 83%

Zero-resource Machine Translation by Multimodal Encoder-decoder Network with Multimedia Pivot

Hideki Nakayama, Noriki Nishida

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Some error corrections in Sect.2.2 and Table 5, Machine Translation, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22709 2026-07-28 cs.CV cs.AI cs.CL 新提交 83%

RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

RMS@CC-MMD 2026:通过几何交互和多视图共识进行多模态厌女症检测

Md. Ajwad Hossain

机构 * Chittagong University of Engineering & Technology (CUET)(吉大港工程技术大学)

专题命中 多模态训练与对齐 :multimodal(title,comments);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对多模态厌女症检测问题,提出GeoMVC方法,通过几何交互层建模跨模态对齐,用多视图共识策略减轻分布偏移,在相关挑战中取得一定排名成绩,凸显特定文化背景下建模的挑战。

Comments Accepted for the CC-MMD Grand Challenge at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17372 2024-10-08 cs.IR 83%

An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders

Youhua Li, Hanwen Du, Yongxin Ni, Yuanqi He, Junchen Fu, Xiangyan Liu, Qi Guo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

Comments An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00102 2023-04-10 cs.CV cs.AI cs.MM 83%

Dynamic Multimodal Fusion

Zihui Xue, Radu Marculescu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM;multi-modal(comments)

Comments Accepted by 6th Multi-Modal Learning and Applications Workshop (MULA), CVPR 2023. Code available at: https://github.com/zihuixue/DynMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17444 2022-12-06 cs.LG 83%

Multimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal Representations

Sijie Mai, Ying Zeng, Haifeng Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments This paper is accepted by IEEE Transactions on Multimedia. This version addresses some mistakes and typos in the original paper. The appendix is available at https://github.com/TmacMai/Multimodal-Information-Bottleneck/blob/main/appendix.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07275 2018-08-23 cs.AI cs.CV cs.MM 83%

CentralNet: a Multilayer Approach for Multimodal Fusion

Valentin Vielzeuf, Alexis Lechervy, Stéphane Pateux, Frédéric Jurie

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Journal ref European Conference on Computer Vision Workshops: Multimodal Learning and Applications, Sep 2018, Munich, Germany. https://mula2018.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11171 2026-06-19 cs.CV cs.AI 版本更新 82%

TerraMind: Large-Scale Generative Multimodality for Earth Observation

TerraMind:面向地球观测的大规模生成式多模态模型

Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Longépé

机构 * IBM Research – Europe(IBM欧洲研究院) ETH Zurich(苏黎世联邦理工学院) Forschungszentrum Jülich(尤利希研究中心) European Space Agency(欧洲航天局) Φ \Phi -Lab(Φ实验室) NASA IMPACT University of Iceland(爱沙尼亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);multimodal foundation model(abstract)

AI总结 提出首个任意到任意生成式多模态基础模型TerraMind,通过双尺度表示(token级和像素级)预训练,实现零样本/少样本应用,并引入“模态思考”能力,在PANGAEA等基准上达到领先性能。

Comments Accepted at ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18681 2025-07-04 cs.CL cs.AI 82%

Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models

Yazhou Zhang, Chunwang Zou, Bo Wang, Jing Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI;multimodal(comments)

Comments Our original goal was to use Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection (arXiv:2506.19420) to replace Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models (arXiv:2503.18681). Due to various reasons, both versions were released, so we would like to withdraw the latter

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16201 2026-08-18 cs.LG 新提交 82%

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

面向基于大语言模型(LLM)的多模态情感分析的多粒度情感集成

Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao

机构 * Fuzhou University(福州大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Jiangxia University(江夏大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。

Comments Accepted to NLPCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09572 2026-08-11 cs.LG 新提交 82%

Hyperbolic Multimodal Continual Learning

双曲多模态持续学习

Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本研究针对双曲多模态持续学习的遗忘问题,建立理论基础推导了保留几何结构的持续学习框架,经实验验证其有效性。

Comments ICML 2026. 33 pages, 11 figures. Code: ICML" target="_blank" rel="noopener">https://github.com/HUBERILT/HMCL_ICML

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07895 2026-08-11 cs.RO cs.LG 新提交 82%

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

多模态机器人演示中指令-轨迹不匹配的审计

Simon Holk, Ryosuke Takanami, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Yueh-Hua Wu, Kei Ota

机构 * AI Robot Association (AIRoA)(人工智能机器人协会(AIRoA)) The University of Tokyo(东京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 针对多模态机器人演示中指令-轨迹不匹配问题,提出无需训练的MMPF审计框架,在LIBERO基准及真实机器人数据上实现最优ITM检测与标签修正,可提升下游策略学习性能并展示过滤演示的权衡。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L). 8 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06211 2026-08-11 cs.CL cs.AI eess.AS 版本更新 82%

LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

LF²AR:考虑分层动态以改进语言模型的多模态适配

Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer

机构 * Université de Toulon, Aix-Marseille Université, CNRS, LIS, France(法国图卢兹大学、马赛大学、CNRS、LIS) Department of Engineering, University of Cambridge, UK(剑桥大学工程系) Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA(三菱电机研究实验室(MERL)) LIA, Avignon Université, France(法国阿维尼翁大学LIA) LIUM, Le Mans Université, France(法国勒芒大学LIUM) Zenidoc, Marseille, France(法国马赛Zenidoc)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本研究提出LF²AR架构,通过分层抽象-细化动态设计适配机制,在文本转图像、语音模态上提升语言模型性能,支持1.9倍生成加速。

Comments Published as a conference paper at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19597 2026-08-10 cs.LG stat.ML 82%

The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence

对比表示学习的几何力学:对齐势、熵分散和跨模态散度

Yichao Cai, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

AI总结 本文通过测度论框架,在大批量极限下证明InfoNCE目标与确定性能量景观的等价性,揭示单模态与对称多模态之间的几何分岔,并指出跨模态散度项导致模态间隙。

Comments 54 Pages, ICML 2026 (Refined document aesthetics for clearer reading)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04452 2026-08-06 cs.CV cs.AI cs.CL 新提交 82%

Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

Q-CueGraph:用于多模态推理的查询条件化视觉证据图

Pengcheng Pan, Xinfang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Q-CueGraph是一种用于多模态推理的查询条件化视觉证据图,它为冻结读取器生成受预算约束的坐标级观测,在多个基准测试中显著提升了推理性能,尤其适用于证据可定位、问题能区分位置且分辨率受限的场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04234 2026-08-06 math.ST cs.LG stat.ML stat.TH 新提交 82%

Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport

基于联合核熵 gromov-wasserstein 最优传输的多模态对齐

Yixuan Florence Wu, Yilun Zhu, Naichen Shi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 针对跨模态配对数据稀缺的场景,提出 JK-EGW 框架实现多模态对齐,理论样本复杂度匹配标准最优传输,实验中在数据稀缺的预训练编码器嵌入对齐任务上性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02769 2026-08-05 stat.ME cs.LG math.ST stat.ML stat.TH 新提交 82%

DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing

DAIF:一种基于数据的近似消息传递多模态监督学习中间融合框架

Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本研究提出DAIF数据自适应中间融合框架,结合随机矩阵理论与非参数依赖度量,通过近似消息传递生成去噪特征,在模拟及两个多模态数据集上的预测任务中表现优于或媲美现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01282 2026-08-04 stat.ME 新提交 82%

Multimodal domain adaptation under label shift and blockwise missing modalities

标签偏移与分块模态缺失下的多模态域适应

Zebin Wang, Ziang Dou, Molei Liu, Tianxi Cai

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 针对标签偏移与分块模态缺失的多模态域适应问题,提出参考锚定方法,结合代理标签辅助策略,在模拟与RCC应用中实现了分布偏移下的稳定预测。

Comments 49 pages, 3 figures, 15 tables; includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29213 2026-08-03 cs.IR cs.LG 新提交 82%

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

GALA:淘宝商购推荐系统中用于自适应多模态表示的生成对齐学习

Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao, Jia Jia

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract)

AI总结 本文针对外卖推荐系统多模态融合难、语义与行为对齐不足的问题,提出三阶段流程GALA,通过生成式RL对齐阶段弥合预训练-微调差距,在淘宝商购部署后提升了订单量与AUC等指标。

Comments 13 pages, 12 figures, 5 tables. Accepted at the 2026 IEEE International Conference on Data Engineering (ICDE 2026), Industry and Applications Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27289 2026-07-31 cs.LG 新提交 82%

TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification

TIER-MoE:用于生物医学分类多模态融合的、基于条件模态风险的信任感知专家路由

Yu Chang, Anzhe Cheng, Chenwei Wu, Zhuoran Wang, Jiahao Chen, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Paul M. Thompson, Liyue Shen, Paul Bogdan

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 该研究提出TIER-MoE风险引导子空间混合专家模型,用于生物医学分类多模态融合,可提升预测性能与概率校准,在多数据集上优于现有最优方法且具备强零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27260 2026-07-31 cs.LG 新提交 82%

Regularizing modality contribution drift in multimodal continual learning

多模态持续学习中模态贡献漂移的正则化

Zhen Zhang, Jielei Chu, Bin Liu, Tianrui Li

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 针对多模态持续学习中的模态贡献漂移问题,提出含基于重放和无重放版本的CMCDR方法,经实验验证其通用性与有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15928 2026-07-28 cs.LG 版本更新 82%

Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment

通过标签条件对比对齐实现成人到儿科心电图转换的知识引导跨模态融合

Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu

机构 * School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院) Nanjing Medical University(南京医科大学) Zhengzhou University(郑州大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

AI总结 研究针对成人与儿科心电图转换问题,提出知识引导的跨模态融合框架PEACE,通过标签条件对比对齐等方法,在有限监督下实现更好的儿科心电图解释,消融实验证明标签条件知识对齐是关键驱动因素。

Comments This article was accidentally submitted as a new arXiv paper instead of a replacement of arXiv:2605.00647. Please refer to arXiv:2605.00647 for the correct and updated version

详情

展开后加载摘要…

URL PDF HTML 收藏