arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-09 至 2026-03-09 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 13 篇

2603.06508 2026-03-09 cs.LG 82%

When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models

当一种模态统治一切:多模态扩散模型中的后门模态崩溃

Qitong Wang, Haoran Dai, Haotian Zhang, Christopher Rasmussen, Binghui Wang

机构 * University of Delaware(德克萨斯大学) Illinois Institute of Technology(伊利诺伊理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文研究了多模态扩散模型中后门攻击的模态崩溃现象,发现攻击常依赖单一模态,跨模态交互反而降低攻击效果,揭示了现有评估的盲点。

Comments Accepted to the ICLR 2026 Workshop on Principled Design for Trustworthy AI. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06544 2026-03-09 cs.CV 79%

Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving

多源多模态数据中冗余建模与测量用于自动驾驶

Yuhan Zhou, Mehri Sattari, Haihua Chen, Kewei Sha

机构 * Dept. of Information Science University of North Texas Denton, Texas, USA(信息科学系 诺克斯维尔大学 德顿, 德克萨斯州, 美国) Dept. of Data Science University of North Texas Denton, Texas, USA(数据科学系 诺克斯维尔大学 德顿, 德克萨斯州, 美国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文研究了自动驾驶中多源多模态数据的冗余问题,通过实验展示了删除冗余标签对目标检测性能的提升作用,并揭示了数据质量与性能之间的直接联系。

Comments This paper has been accepted by the Fourth IEEE International Conference on Mobility: Operations, Services, and Technologies (MOST) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06507 2026-03-09 cs.CV 79%

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

自监督流匹配用于可扩展的多模态合成

Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach

机构 * Black Forest Labs(黑森林实验室) MIT(麻省理工学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 Self-Flow通过自监督流匹配方法,在无需外部监督的情况下提升多模态生成的性能和扩展性。

Comments project webpage: https://bfl.ai/research/self-flow

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06147 2026-03-09 cs.CV 79%

Longitudinal NSCLC Treatment Progression via Multimodal Generative Models

多模态生成模型用于非小细胞肺癌治疗进展的纵向预测

Massimiliano Mantegna, Elena Mulero Ayllón, Alice Natalina Caragliano, Francesco Di Feola, Claudia Tacconi, Michele Fiore, Edy Ippolito, Carlo Greco, Sara Ramella, Philippe C. Cattin, Paolo Soda, Matteo Tortora, Valerio Guarrasi

机构 * Unit of Artificial Intelligence and Computer Systems, Department of Engineering, Università Campus Bio-Medico di Roma, Italy(人工智能与计算机系统单位,工程系,罗马大学生物医学学院) Multi-Specialist Clinical Institute for Orthopaedic Trauma Care (COT), Messina, Italy(骨科创伤护理多学科临床研究所(COT),意大利Messina) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering, Umeå University, Sweden(诊断与介入系,放射物理,生物医学工程,乌梅大学,瑞典) Operative Research Unit of Radiation Oncology, Fondazione Policlinico Universitario Campus Bio-Medico, Rome, Italy(放射肿瘤手术研究单位,大学生物医学学院基金会,罗马,意大利) Research Unit of Radiation Oncology, Department of Medicine and Surgery, Università Campus Bio-Medico di Roma, Italy(放射肿瘤研究单位,医学与外科系,罗马大学生物医学学院,意大利) Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(生物医学工程系,巴塞尔大学,瑞士Allschwil) Department of Naval, Electrical, Electronics and Telecommunications Engineering, University of Genoa, Italy(海军、电气、电子与电信工程系,热那亚大学,意大利)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出虚拟治疗框架,利用多模态生成模型预测NSCLC治疗进展,验证扩散模型在生成稳定肿瘤演变轨迹方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06043 2026-03-09 cs.CV 79%

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

通过理解学习生成:基于理解的内在奖励用于统一多模态模型

Jiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang, Shuohuan Wang, Yu Sun, Hua Wu, Qingming Huang, Haifeng Wang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Peking University(北京大学) University of Chinese Academy of Sciences(中国科学院大学) Sun Yat-sen University(中山大学) Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出GvU机制,通过内在奖励提升统一多模态模型的生成能力,同时增强其视觉理解能力,缩小理解与生成之间的能力差距。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05800 2026-03-09 cs.DC cs.AI 79%

StreamWise: Serving Multi-Modal Generation in Real-Time at Scale

StreamWise: 实时大规模多模态生成服务

Haoran Qiu, Gohar Irfan Chaudhry, Chaojie Zhang, Íñigo Goiri, Esha Choukse, Rodrigo Fonseca, Ricardo Bianchini

机构 * Microsoft Azure Research(微软Azure研究院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

AI总结 StreamWise通过实时播客视频生成,整合LLM、文本转语音和视频音频生成,实现高效多模态实时服务,兼顾延迟、成本和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06057 2026-03-09 cs.CV cs.AI cs.LG cs.SD 62%

TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation

TempoSyncDiff: 压缩时间一致扩散用于低延迟音频驱动说话头生成

Soumya Mazumdar, Vineet Kumar Rakesh

机构 * Computer and Informatics Group, Variable Energy Cyclotron Centre(计算机与信息组,变能循环中心)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 TempoSyncDiff通过压缩扩散模型实现低延迟音频驱动说话头生成,结合教师-学生压缩和时间正则化以提高生成稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13687 2026-03-09 cs.CV 57%

Towards Scalable Pre-training of Visual Tokenizers for Generation

面向生成任务的视觉分词器可扩展预训练

Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) MiniMax

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

AI总结 VTP通过联合优化图像-文本对比、自监督和重建损失,提升视觉分词器的生成性能和扩展性。

Comments Our pre-trained models are available at https://github.com/MiniMax-AI/VTP

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06038 2026-03-09 cs.CV cs.GR 57%

FontUse: A Data-Centric Approach to Style- and Use-Case-Conditioned In-Image Typography

FontUse: 一种以数据为中心的方法用于风格和使用场景条件下的图像内字体设计

Xia Xin, Yuki Endo, Yoshihiro Kanamori

机构 * University of Tsukuba(茨口大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 FontUse提出了一种以数据为中心的方法,通过构建大规模字体数据集并结合多模态模型,提升文本生成中字体风格和使用场景的控制能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06032 2026-03-09 cs.CV 57%

StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision

StruVis: 通过结构化视觉进行推理的文本到图像生成增强

Yuanhuiyi Lyu, Kaiyu Lei, Ziqiao Weng, Xu Zheng, Lutao Jiang, Teng Li, Yangfu Li, Ziyuan Huang, Linfeng Zhang, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Ant Group(蚂蚁集团) Shanghai Jiao Tong University(上海交通大学) Hong Kong University of Science and Technology(香港科技大学) East China Normal University(华东师范大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 StruVis通过结构化视觉引导推理,提升基于推理的文本到图像生成性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06014 2026-03-09 cs.CV 57%

EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation

EffectMaker:统一推理与生成以实现定制化视觉效果创建

Shiyuan Yang, Ruihuang Li, Jiale Tao, Shuai Shao, Qinglin Lu, Jing Liao

机构 * Tencent Hunyuan(腾讯文言) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 EffectMaker通过统一推理生成框架实现定制化视觉效果创建,利用多模态模型和扩散变换器提升效果质量和一致性,构建大规模数据集增强泛化能力。

Comments Project page: https://effectmaker.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05769 2026-03-09 cs.CV 57%

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

逐层实例绑定用于文本到图像扩散变换器中的区域和遮挡控制

Ruidong Chen, Yancheng Bai, Xuanpu Zhang, Jianhao Zeng, Lanjun Wang, Dan Song, Lei Sun, Xiangxiang Chu, Anan Liu

机构 * Tianjin University(天津大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 LayerBind通过逐层实例绑定实现文本到图像扩散变换器中的区域和遮挡控制,无需训练即可提升生成质量和遮挡管理能力。

Comments Accepted by CVPR26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05518 2026-03-09 cs.HC cs.CV 57%

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

CoEditor++:基于指令的视觉编辑 via 认知推理

Minheng Ni, Yutao Fan, Zhengyuan Yang, Yeli Shen, Yuxiang Wei, Yaowen Zhang, Lijuan Wang, Lei Zhang, Wangmeng Zuo

机构 * Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology(哈尔滨工业大学) Microsoft(微软公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 CoEditor++通过认知推理实现基于指令的视觉编辑,无需训练即可在通用和负责任的编辑任务中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏