arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-04 至 2026-03-04 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2603.03276 2026-03-04 cs.CV 83%

Beyond Language Modeling: An Exploration of Multimodal Pretraining

超越语言模型:多模态预训练的探索

Shengbang Tong, David Fan, John Nguyen, Ellis Brown, Gaoyue Zhou, Shengyi Qian, Boyang Zheng, Théophane Vallaeys, Junlin Han, Rob Fergus, Naila Murray, Marjan Ghazvininejad, Mike Lewis, Nicolas Ballas, Amir Bar, Michael Rabbat, Jakob Verbeek, Luke Zettlemoyer, Koustuv Sinha, Yann LeCun, Saining Xie

机构 * FAIR, Meta(FAIR、Meta) New York University(纽约大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文通过多模态预训练探索,揭示了视觉与语言数据的互补性及统一预训练对世界建模的促进作用,并提出MoE架构解决多模态扩展的不对称性问题。

Comments Project website at https://beyond-llms.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02532 2026-03-04 cs.CV 83%

EIMC: Efficient Instance-aware Multi-modal Collaborative Perception

EIMC: 高效实例感知多模态协作感知

Kang Yang, Peng Wang, Lantao Li, Tianci Bu, Chen Sun, Deying Li, Yongcai Wang

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Sony Research and Development Center China(索尼(中国)研发有限公司) National University of Defense Technology(国防科技大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 EIMC通过实例感知的多模态协作感知方法,提升自动驾驶安全性,减少带宽使用,实现高效且准确的3D感知。

Comments 9 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02505 2026-03-04 cs.CV 83%

SGMA: Semantic-Guided Modality-Aware Segmentation for Remote Sensing with Incomplete Multimodal Data

SGMA: 基于不完整多模态数据的语义引导模态感知遥感语义分割

Lekang Wen, Liang Liao, Jing Xiao, Mi Wang

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(武汉大学测绘遥感信息工程国家重点实验室) Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州研究院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 SGMA通过语义引导和模态感知方法解决不完整多模态数据下的语义分割问题,提升模态间平衡性和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02474 2026-03-04 cs.IR cs.AI 83%

Q-BERT4Rec: Quantized Semantic-ID Representation Learning for Multimodal Recommendation

Q-BERT4Rec: 量化语义-ID表示学习用于多模态推荐

Haofeng Huang, Ling Gai

机构 * University of Shanghai for Science and Technology(上海科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 Q-BERT4Rec通过统一语义表示与量化建模,提升多模态推荐的性能和可解释性。

Comments Submitted to KDD2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05612 2026-03-04 cs.LG cs.AI 83%

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Shuffle-R1: 通过数据导向的动态洗牌提升多模态大语言模型的强化学习框架

Linghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Bin Qin, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(MiLM Plus,小米公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 Shuffle-R1通过动态洗牌和轨迹采样提升多模态大语言模型的强化学习效率,实现更高效的训练效果。

Comments This paper has been accepted by ICLR 2026. Conference link: https://iclr.cc/virtual/2026/poster/10007559 OpenReview link: https://openreview.net/forum?id=mYP33u1QBK Project page at: https://xenozlh.github.io/Shuffle-R1/

Journal ref The Fourteenth International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01592 2026-03-04 cs.CV cs.AI 81%

DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter

DMTrack: 基于双适配器的时空多模态跟踪

Weihong Li, Shaohua Dong, Haonan Lu, Yanhao Zhang, Heng Fan, Libo Zhang

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Department of Computer Science and Engineering, University of North Texas(北卡罗来纳州立大学计算机科学与工程系) OPPO AI Center(OPPO人工智能中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 DMTrack通过双适配器架构实现时空多模态跟踪,利用STMA和PMCA模块提升跨模态融合效果,达到先进性能。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02629 2026-03-04 cs.CV 79%

Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective

迈向增量统一多模态异常检测:从信息瓶颈视角增强多模态去噪

Kaifang Long, Lianbo Ma, Jiaqi Liu, Liming Liu, Guoyang Xie

机构 * Software College, Northeastern University, China(东北大学软件学院) CATL, China(宁德时代)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出IB-IUMAD框架,通过Mamba解码器和信息瓶颈融合模块解决多模态异常检测中的灾难性遗忘问题,提升模型对新兴对象的适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02672 2026-03-04 physics.plasm-ph 78%

PanoMHD: Multimodal Modelling of Plasma Dynamics towards Tokamak Control

PanoMHD:面向托卡马克控制的等离子体动力学多模态建模

Hyeongjun Noh, Chweeho Heo, Xiaotian Gao, Yong-Su Na

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 PanoMHD通过多模态建模提升等离子体稳定性预测,实现对托卡马克控制的先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22903 2026-03-04 cs.IR cs.LG 78%

PSQE: A Theoretical-Practical Approach to Pseudo Seed Quality Enhancement for Unsupervised Multimodal Entity Alignment

PSQE:一种伪种子质量增强的理论-实践方法用于无监督多模态实体对齐

Yunpeng Hong, Chenyang Bu, Jie Zhang, Yi He, Di Wu, Xindong Wu

机构 * Key Laboratory of Knowledge Engineering with Big Data (the Ministry of Education of China), Hefei University of Technology(大数据知识工程重点实验室(教育部)、合肥工业大学) Department of Data Science, College of William and Mary(威廉与玛丽学院数据科学系) College of Computer and Information Science, Southwest University(西南大学计算机与信息科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 PSQE通过多模态信息和聚类重采样提升伪种子质量,改善无监督多模态实体对齐的精度与图覆盖平衡性。

Comments 2026 SIGKDD Accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11062 2026-03-04 cs.LG cs.IR 78%

MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start Recommendation

MoToRec:基于稀疏正则化的多模态分词用于冷启动推荐

Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 MoToRec通过稀疏正则化的多模态分词方法,解决冷启动推荐中的数据稀疏性和新物品表示问题,提升推荐系统性能。

Comments Accepted to AAAI 2026 (Main Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02250 2026-03-04 cs.SD eess.AS 74%

SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models

SGPA: 基于频谱的语音对齐用于多模态大语言模型中可行的谢普利值解释

Paweł Pozorski, Jakub Muszyński, Maria Ganzha

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 eess.AS

AI总结 SGPA通过结合连接主义时间分类和频谱边界细化,实现了多模态大语言模型中可行的音频解释,显著减少了模型评估次数并保持了全局轮廓。

Comments Submitted for admission in Interspeech 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02560 2026-03-04 cs.CV 70%

CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration

CAWM-Mamba:一种用于红外可见图像融合和复合恶劣天气恢复的统一模型

Huichun Liu, Xiaosong Li, Zhuangfan Huang, Tao Ye, Yang Liu, Haishu Tan

机构 * School of Physics(物理学院) Optoelectronic Engineering, Foshan University, Foshan 528225, China(光学电子工程学院,佛山大学,佛山528225,中国) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology, Foshan 528225, China(粤港澳联合智能微纳光电子技术实验室,佛山528225,中国) School of Mechanical Electronic(机械电子学院) Information Engineering, China University of Mining and Technology, Beijing 100083, China(信息工程学院,中国矿业大学,北京100083,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 CAWM-Mamba提出了一种统一模型,用于红外可见图像融合和复合恶劣天气恢复,通过三个关键模块提升多退化场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14337 2026-03-04 eess.IV cs.CV 70%

Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition

无监督可变形图像配准与局部-全局注意力及图像分解

Zhengyong Huang, Xingwen Sun, Xuting Chang, Ning Jiang, Yao Wang, Jianfei Sun, Hongbin Han, Yao Sui

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学医学部医学技术研究所) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Department of Radiology, Peking University Third Hospital(北京大学第三医院放射科) Department of Pediatrics, Peking University First Hospital(北京大学第一医院儿科) Pediatric Epilepsy Center, Peking University First Hospital(北京大学第一医院儿童癫痫中心) School of Biological Science and Medical Engineering, Southeast University(东南大学生物科学与医学工程学院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出LGANet++,通过局部-全局注意力机制和图像分解技术,实现更准确、鲁棒的无监督可变形图像配准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23652 2026-03-04 cs.CV cs.AI 62%

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

面向MRI多器官异常检测的3D模态感知预训练

Haowen Zhu, Ning Yin, Xiaogen Zhou

机构 * School of Electronic, Electrical Engineering and Physics, Fujian University of Technology(福建工程学院电子电气工程学院) School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Department of Medical Imaging, Suzhou Traditional Chinese Medicine Hospital, China(苏州中医医院影像科)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MedMAP框架,通过3D MRI的模态感知预训练提升多器官异常检测性能,实验表明其优于现有视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02924 2026-03-04 cs.CV 57%

HDINO: A Concise and Efficient Open-Vocabulary Detector

HDINO:一种简洁且高效的开放词汇检测器

Hao Zhang, Yiqun Wang, Qinran Lin, Runze Fan, Yong Li

机构 * Chongqing University(重庆大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 HDINO提出了一种简洁高效的开放词汇检测方法,通过两阶段训练策略和轻量级特征融合模块,在无需人工数据编纂的情况下实现了优于现有方法的检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02522 2026-03-04 cs.CV 57%

NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining

NeighborMAE: 利用邻近遥感图像间的空间依赖性在掩码自编码器预训练中进行建模

Liang Zeng, Valerio Marsocci, Wufan Zhao, Andrea Nascetti, Maarten Vergauwen

机构 * KU Leuven(卢森堡大学) ESA Φ \Phi -lab(欧洲航天局Φ实验室) HKUST(GZ)(香港科技大学(珠海)) KTH(皇家理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 NeighborMAE通过联合重建邻近遥感图像,利用空间依赖性提升掩码自编码器在遥感图像预训练中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏