arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2512.02088 2025-12-03 eess.IV cs.AI cs.CV cs.LG 81%

Comparing Baseline and Day-1 Diffusion MRI Using Multimodal Deep Embeddings for Stroke Outcome Prediction

比较基线和第1天扩散磁共振成像用于中风预后预测的多模态深度嵌入

Sina Raeisadigh, Myles Joshua Toledo Tan, Henning Müller, Abderrahmane Hedjoudje

机构 * 1 Department of Computer Science, University of Geneva, Switzerland 2 Department of Electrical \& Computer Engineering, University of Florida, FL, USA 3 Service of Medical Informatics, University Hospital of Geneva, Switzerland 4 Department of Imaging Medical Informatics, University of Geneva, Switzerland

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究通过多模态深度嵌入方法,利用基线和治疗后1天的扩散MRI数据,结合临床特征和病变体积,预测急性缺血性中风患者3个月的功能预后。

Comments 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12170 2025-12-02 cs.CV cs.AI 81%

Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective

重新思考多模态点云补全:一种补全-修正视角

Wang Luo, Di Wu, Hengyuan Na, Yinlin Zhu, Miao Hu, Guocong Quan

机构 * Wang Luo, Di Wu, Hengyuan Na, Yinlin Zhu, Miao Hu, Guocong Quan(作者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PGNet通过补全-修正范式,结合多阶段框架和双特征编码,实现更鲁棒的点云补全,提升重建精度与结构一致性。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00677 2025-12-02 cs.CV cs.AI 81%

Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer

Dynamic-eDiTor: 基于文本的无训练4D场景编辑与多模态扩散变换器

Dong In Lee, Hyungjun Doh, Seunggeun Chi, Runlin Duan, Sangpil Kim, Karthik Ramani

机构 * Purdue University(普渡大学) Korea University(韩国大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 Dynamic-eDiTor通过多模态扩散变换器和4DGS实现无训练文本驱动的4D场景编辑,提升多视图和时间一致性。

Comments 4D Scene Editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20561 2025-12-02 cs.CV cs.CL 81%

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward

在统一多模态模型中,理解是否会影响生成?从分析到未来路径

Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin, Bin Lin, Zongjian Li, Bin Zhu, Weihao Yu, Li Yuan

机构 * Peking University(北京大学) Chongqing University(重庆大学) HKU MMLab(香港大学多模态实验室) PengCheng Laboratory(鹏城实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 研究通过UniSandbox分析统一多模态模型中理解与生成之间的差距,发现显式链式思维和自训练方法能有效弥合这一差距,并揭示查询架构的潜在CoT特性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10125 2025-12-02 cs.CV cs.MM 81%

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

代理调优:为以主题驱动的图像生成定制多模态自回归模型

Yi Wu, Shengju Qian, Lingting Zhu, Lei Liu, Wandi Qiao, Ziqiang Li, Lequan Yu, Bin Li

机构 * University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出代理调优方法,通过扩散模型增强AR模型在主题驱动图像生成中的能力,揭示了弱到强泛化现象,提升了多主题组合和上下文理解的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18428 2025-11-26 cs.AI cs.CL 81%

Multi-Modal Data Exploration via Language Agents

通过语言代理进行多模态数据探索

Farhad Nooralahzadeh, Yi Zhang, Jonathan Furst, Kurt Stockinger

机构 * Zurich University of Applied Sciences(瑞士应用科学大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出M$^2$EX系统,通过语言代理实现多模态数据探索,优于现有系统在准确性和性能指标上

Comments Accepted to the IJCNLP AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14993 2025-11-26 cs.AI cs.CV 81%

Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification

多模态生成式AI:多模态大语言模型、扩散模型与统一

Xin Wang, Yuwei Zhou, Bin Huang, Hong Chen, Wenwu Zhu

机构 * Department of Computer Science, Beijing Information Science and Technology National Research Center, Tsinghua University(计算机系,北京信息科学与技术国家研究中心,清华大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文综述了多模态生成式AI,涵盖多模态LLMs、扩散模型及统一模型的设计与应用,探讨了统一理解和生成的方法及未来研究方向。

Comments 21 pages, 10 figures, 3 tables

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03227 2025-11-07 cs.HC cs.AI cs.MM 81%

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

Alexander Htet Kyaw, Lenin Ravindranath Sivalingam

机构 * Massachusetts Institute of Technology(麻省理工学院) Microsoft Research(微软研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted to NeurIPS 2025, Conference on Neural Information Processing Systems, Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21448 2025-11-06 eess.AS cs.CV cs.SD 81%

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang, Qian Chen, Zhou Zhao, Wei Xue

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Tongyi Fun Team, Alibaba Group(阿里云团队) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted by NeurIPS 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19755 2025-11-04 cs.LG cs.AI cs.CV 81%

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

Jiacheng Liu, Xinyu Wang, Yuqi Lin, Zhikai Wang, Peiru Wang, Peiliang Cai, Qinming Zhou, Zhengan Yan, Zexuan Yan, Zhengyi Shi, Chang Zou, Yue Ma, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 22 pages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27632 2025-11-03 cs.CV cs.AI 81%

Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation

Riccardo Brioschi, Aleksandr Alekseev, Emanuele Nevali, Berkay Döner, Omar El Malki, Blagoj Mitrevski, Leandro Kieliger, Mark Collier, Andrii Maksai, Jesse Berent, Claudiu Musat, Efi Kokiopoulou

机构 * EPFL(苏黎世联邦理工学院) Google DeepMind(谷歌DeepMind)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 18 figures, GitHub link: https://github.com/google-deepmind/sketch_to_layout, accept at ICCV 2025 Workshop (HiGen)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26105 2025-10-31 cs.CV cs.AI cs.CR 81%

Security Risk of Misalignment between Text and Image in Multi-modal Model

Xiaosen Wang, Zhijin Ge, Shaokang Wang

机构 * Xidian University(西安电子科技大学) Shanghai Jiaotong University(上海交通大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24514 2025-10-29 cs.CV cs.CL 81%

Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

Huanyu Zhang, Wenshan Wu, Chengzu Li, Ning Shang, Yan Xia, Yangyu Huang, Yifan Zhang, Li Dong, Zhang Zhang, Liang Wang, Tieniu Tan, Furu Wei

机构 * MSR(微软研究院) UCAS(中国科学院自动化研究所) CASIA(中国科学院自动化研究所) Cambridge(剑桥大学) NJU(南京大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22684 2025-10-28 cs.CV cs.CL 81%

RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

Jiuniu Wang, Gongjie Zhang, Quanhao Qian, Junlong Gao, Deli Zhao, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22521 2025-10-28 cs.CV cs.AI cs.IR cs.LG 81%

Open Multimodal Retrieval-Augmented Factual Image Generation

Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi, Yupeng Hu, Liqiang Nie

机构 * Shandong University(山东大学) National University of Singapore(新加坡国立大学) Kuaishou Technology(快手科技) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18400 2025-10-22 eess.IV cs.AI cs.CV 81%

A Multimodal Deep Learning Approach for White Matter Shape Prediction in Diffusion MRI Tractography

Yui Lo, Yuqian Chen, Dongnan Liu, Leo Zekelman, Jarrett Rushmore, Yogesh Rathi, Nikos Makris, Alexandra J. Golby, Fan Zhang, Weidong Cai, Lauren J. O'Donnell

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Paper accepted to Human Brain Mapping. 25 pages, 3 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09121 2025-10-21 cs.CV cs.AI 81%

MSDM: Generating Task-Specific Pathology Images with a Multimodal Conditioned Diffusion Model for Cell and Nuclei Segmentation

Dominik Winter, Mai Bui, Monica Azqueta Gavaldon, Nicolas Triltsch, Marco Rosati, Nicolas Brieu

机构 * AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13253 2025-10-20 cs.CV cs.AI cs.LG 81%

End-to-End Multi-Modal Diffusion Mamba

Chunhao Lu, Qiang Lu, Meichen Dong, Jake Luo

机构 * China University of Petroleum-Beijing(中国石油大学(北京)) Leyard Optoelectronic(莱亚德光电) University of Wisconsin-Milwaukee(威斯康星大学密尔沃基分校)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08209 2025-10-17 cs.CV cs.AI cs.LG 81%

Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision

Shengcao Cao, Liang-Yan Gui, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 81%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21787 2025-10-13 cs.CV cs.CL 81%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05661 2025-10-08 cs.CV cs.MM 81%

When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach

Daniel Gonzálbez-Biosca, Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano

机构 * eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya(eHealth中心,计算机科学、多媒体与电信学院,开放大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 81%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17121 2025-10-06 cs.CL cs.AI 81%

NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation

Weiming Wu, Jin Ye, Zi-kang Wang, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26644 2025-10-01 cs.CV cs.AI cs.LG 81%

Stitch: Training-Free Position Control in Multimodal Diffusion Transformers

Jessica Bader, Mateusz Pach, Maria A. Bravo, Serge Belongie, Zeynep Akata

机构 * Technical University of Munich(慕尼黑技术大学) Helmholtz Munich(海德堡-慕尼黑研究所) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Copenhagen(哥本哈根大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21360 2025-09-29 cs.CV cs.AI 81%

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21107 2025-09-26 cs.RO cs.AI cs.CV cs.LG 81%

Cross-Modal Instructions for Robot Motion Generation

William Barron, Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi

机构 * College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学) School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21476 2025-09-23 cs.CV cs.AI 81%

GarmentDiffusion: 3D Garment Sewing Pattern Generation with Multimodal Diffusion Transformers

Xinyu Li, Qi Yao, Yuanda Wang

机构 * Shenfu Research(沈孚研究所) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16197 2025-09-22 cs.CV cs.CL cs.LG 81%

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan, Haotian Zhang, Haoshuo Huang, Bowen Zhang, Jialing Tong, Haoxuan You, Xianzhi Du, Zhe Gan, Hyunjik Kim, Chao Jia, Zhenbang Wang, Yinfei Yang, Mingfei Gao, Zi-Yi Dou, Wenze Hu, Chang Gao, Dongxu Li, Philipp Dufter, Zirui Wang, Guoli Yin, Zhengdong Zhang, Chen Chen, Yang Zhao, Ruoming Pang, Zhifeng Chen

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15553 2025-09-22 cs.CV cs.AI stat.AP 81%

Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification

Tian Lan, Yiming Zheng, Jianxin Yin

机构 * School of Statistics, Renmin University of China(中国人民大学统计学院) Center for Applied Statistics and School of Statistics, Renmin University of China(中国人民大学应用统计中心和统计学院)

专题命中 多模态生成 :cross-modal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏