arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2503.22517 2025-04-02 cs.CL cs.AI cs.CV 82%

Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities

Raman Dutt, Harleen Hanspal, Guoxuan Xia, Petru-Daniel Tudosiu, Alexander Black, Yongxin Yang, Steven McDonagh, Sarah Parisot

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18597 2025-03-27 cs.CV cs.AI cs.MM 82%

DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Minghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu, Zhaoyang Zhang, Yong Zhang, Ying Shan, Xiangyu Yue

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments CVPR 2025; 21 pages, 23 figures, Project page: https://onevfall.github.io/project_page/ditctrl ; GitHub repository: https://github.com/TencentARC/DiTCtrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12214 2025-03-18 cs.LG math.DS 82%

Cross-Modal Diffusion for Biomechanical Dynamical Systems Through Local Manifold Alignment

Sharmita Dey, Sarath Ravindran Nair

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17842 2025-02-26 cs.CV cs.LG cs.MM cs.SD eess.AS 82%

MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19353 2025-02-19 cs.CL cs.AI cs.CV 82%

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu, Lun-Wei Ku, C. Lee Giles, Ting-Hao K. Huang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to TACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15188 2025-02-06 cs.CL cs.AI cs.CV cs.LG 82%

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Weijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang, Xi Victoria Lin, Luke Zettlemoyer, Lili Yu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Name change: LlamaFusion to LMFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17811 2025-01-30 cs.AI cs.CL cs.CV 82%

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Research paper. arXiv admin note: text overlap with arXiv:2410.13848

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20378 2024-12-31 cs.CV cs.MM cs.SD eess.AS 82%

Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Bingliang Li, Fengyu Yang, Yuxin Mao, Qingwen Ye, Hongkai Chen, Yiran Zhong

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments AAAI 2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13945 2024-12-24 cs.SE 82%

How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation

Liwen Wang, Yuanyuan Yuan, Ao Sun, Zongjie Li, Pingchuan Ma, Daoyuan Wu, Shuai Wang

专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16045 2024-12-12 cs.CV cs.AI cs.CL cs.LG 82%

Woodpecker: Hallucination Correction for Multimodal Large Language Models

Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, Enhong Chen

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by Science China Information Sciences (SCIS)

Journal ref SCIENCE CHINA Information Sciences, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02368 2024-12-04 cs.AI cs.CL cs.CV 82%

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Christoph Leiter, Simone Paolo Ponzetto, Fahimeh Moafian, Zhixue Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13848 2024-10-18 cs.CV cs.AI cs.CL 82%

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07157 2024-10-10 cs.AI cs.CL cs.CV cs.LG cs.SI 82%

InstructG2I: Synthesizing Images from Multimodal Attributed Graphs

Bowen Jin, Ziqi Pang, Bingjun Guo, Yu-Xiong Wang, Jiaxuan You, Jiawei Han

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 16 pages

Journal ref NeurIPs 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05964 2024-09-17 cs.MM cs.AI cs.CL 82%

Interpretable Multimodal Misinformation Detection with Logic Reasoning

Hui Liu, Wenya Wang, Haoliang Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted by Findings of ACL 23. 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09787 2024-08-20 cs.CL cs.CV cs.MM 82%

Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation

Yunxin Li, Haoyuan Shi, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao, Min Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by SIGGRAPH Asia 2024, Project and Codes: https://github.com/HITsz-TMG/Anim-Director

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15335 2024-07-23 eess.SP 82%

Addressing Out-of-Distribution Challenges in Image Semantic Communication Systems with Multi-modal Large Language Models

Feifan Zhang, Yuyang Du, Kexin Chen, Yulin Shao, Soung Chang Liew

专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11781 2024-06-18 cs.IR 82%

DiffMM: Multi-Modal Diffusion Model for Recommendation

Yangqin Jiang, Lianghao Xia, Wei Wei, Da Luo, Kangyi Lin, Chao Huang

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03776 2024-06-10 cs.CL cs.AI cs.CV cs.IR 82%

XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags

Faisal Tareque Shohan, Mir Tafseer Nayeem, Samsul Islam, Abu Ubaida Akash, Shafiq Joty

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2024 camera ready. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16136 2024-05-28 cs.AI cs.CL cs.LG cs.SD eess.AS 82%

C3LLM: Conditional Multimodal Content Generation Using Large Language Models

Zixuan Wang, Qinkai Duan, Yu-Wing Tai, Chi-Keung Tang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00500 2024-05-22 cs.CV cs.AI cs.MM 82%

Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images

Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi, Alberto Del Bimbo, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ACM Transactions on Multimedia Computing, Communications and Applications (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00923 2024-05-21 cs.CL cs.AI cs.CV 82%

Multimodal Chain-of-Thought Reasoning in Language Models

Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, Alex Smola

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08949 2024-05-20 cs.AI cs.CL cs.CV 82%

EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs

Xiangyu Zhao, Bo Liu, Qijiong Liu, Guangyuan Shi, Xiao-Ming Wu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2024, main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03162 2024-05-07 cs.CV cs.AI cs.CL cs.LG 82%

Advancing Multimodal Medical Capabilities of Gemini

Lin Yang, Shawn Xu, Andrew Sellergren, Timo Kohlberger, Yuchen Zhou, Ira Ktena, Atilla Kiraly, Faruk Ahmed, Farhad Hormozdiari, Tiam Jaroensri, Eric Wang, Ellery Wulczyn, Fayaz Jamil, Theo Guidroz, Chuck Lau, Siyuan Qiao, Yun Liu, Akshay Goel, Kendall Park, Arnav Agharwal, Nick George, Yang Wang, Ryutaro Tanno, David G. T. Barrett, Wei-Hung Weng, S. Sara Mahdavi, Khaled Saab, Tao Tu, Sreenivasa Raju Kalidindi, Mozziyar Etemadi, Jorge Cuadros, Gregory Sorensen, Yossi Matias, Katherine Chou, Greg Corrado, Joelle Barral, Shravya Shetty, David Fleet, S. M. Ali Eslami, Daniel Tse, Shruthi Prabhakara, Cory McLean, Dave Steiner, Rory Pilgrim, Christopher Kelly, Shekoofeh Azizi, Daniel Golden

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13668 2024-04-29 cs.CL cs.AI cs.CV 82%

MAIRA-1: A specialised large multimodal model for radiology report generation

Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Mercy Ranjit, Anton Schwaighofer, Fernando Pérez-García, Valentina Salvatelli, Shaury Srivastav, Anja Thieme, Noel Codella, Matthew P. Lungren, Maria Teodora Wetscherek, Ozan Oktay, Javier Alvarez-Valle

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 18 pages, 9 tables, 5 figures. v2 adds test IDs and image encoder citation. v3 fixes error in NPV/specificity

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14368 2024-04-23 cs.CV cs.AI cs.CL 82%

Graphic Design with Large Multimodal Model

Yutao Cheng, Zhao Zhang, Maoke Yang, Hui Nie, Chunyuan Li, Xinglong Wu, Jie Shao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01858 2024-04-19 cs.LG cs.AI cs.CL cs.CV 82%

Explaining latent representations of generative models with large multimodal models

Mengdan Zhu, Zhenke Liu, Bo Pan, Abhinav Angirekula, Liang Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14652 2024-03-25 cs.CY cs.AI cs.CL cs.MM 82%

MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation

Han Wang, Roy Ka-Wei Lee

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 8 pages, 7 figures, ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03040 2024-02-06 cs.CV cs.AI cs.LG cs.MM 82%

InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions

Yiyuan Zhang, Yuhao Kang, Zhixin Zhang, Xiaohan Ding, Sanyuan Zhao, Xiangyu Yue

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Code, models, and demo are available at https://github.com/invictus717/InteractiveVideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02369 2024-02-06 cs.CV cs.CL cs.MM 82%

M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing

Mohammadreza Mofayezi, Reza Alipour, Mohammad Ali Kakavand, Ehsaneddin Asgari

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01504 2023-12-05 cs.CV cs.AI cs.CL cs.LG 82%

Effectively Fine-tune to Improve Large Multimodal Models for Radiology Report Generation

Yuzhe Lu, Sungmin Hong, Yash Shah, Panpan Xu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Deep Generative Models for Health Workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏