arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2410.13175 2025-05-20 cs.LG cs.AI physics.ao-ph 79%

TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness

Cheng Huang, Pan Mu, Cong Bai, Peter AG Watson

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Camera-ready version. This paper has been accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08803 2025-05-15 cs.LG cs.AI 79%

Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models

Zizhao Hu, Mohammad Rostami, Jesse Thomason

机构 * University of Southern California(南加州大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06303 2025-05-13 cs.LG cs.AI 79%

Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction

Li Yuan, Yi Cai, Xudong Shen, Qing Li, Qingbao Huang, Zikun Deng, Tao Wang

机构 * School of Software Engineering, South China University of Technology, Guangzhou, China(华南理工大学软件学院) Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China(大数据与智能机器人重点实验室) Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学计算机系) School of Electrical Engineering, Guangxi University, Nanning, China(广西大学电气工程学院) Department of Biostatistics & Health Informatics, King’s College London, London, United Kingdom(伦敦国王学院生物统计与健康信息学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04650 2025-05-09 cs.GR cs.AI cs.IR cs.LG 79%

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

Kapil Wanaskar, Gaytri Jena, Magdalini Eirinaki

机构 * Computer Engineering Dept. San José State University(计算机工程系圣何塞州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04996 2025-05-09 cs.CL 79%

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Weixin Liang, Lili Yu, Liang Luo, Srinivasan Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Xi Victoria Lin

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to TMLR 2025; 48 pages

Journal ref Transactions on Machine Learning Research (2025), ISSN: 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21718 2025-05-08 cs.CV 79%

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction

Shiying Li, Xingqun Qi, Bingkun Yang, Chen Weile, Zezhao Tian, Muyi Sun, Qifeng Liu, Man Zhang, Zhenan Sun

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong University of Science and Technology(香港科技大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19341 2025-04-29 cs.RO cs.AI 79%

PolyTouch: A Robust Multi-Modal Tactile Sensor for Contact-rich Manipulation Using Tactile-Diffusion Policies

Jialiang Zhao, Naveen Kuppuswamy, Siyuan Feng, Benjamin Burchfiel, Edward Adelson

机构 * MIT CSAIL

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Nominated for the best paper award at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14320 2025-04-23 cs.HC cs.AI 79%

Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces

Nimisha Karnatak, Adrien Baranes, Rob Marchant, Huinan Zeng, Tríona Butler, Kristen Olson

机构 * University of Oxford(牛津大学) Google DeepMind(谷歌DeepMind) King’s College London(伦敦大学学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at CHI'25 Workshop on Designing and Developing User Interfaces with AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21979 2025-04-23 cs.CV 79%

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Zhonghua Wu, Qingyi Tao, Wentao Liu, Wei Li, Chen Change Loy

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory Research(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) SenseTime Research and Tetras.AI(SenseTime研究部和Tetras.AI) SenseTime Research(商汤科技研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03173 2025-04-23 cs.CV 79%

MObI: Multimodal Object Inpainting Using Diffusion Models

Alexandru Buburuzan, Anuj Sharma, John Redford, Puneet K. Dokania, Romain Mueller

机构 * FiveAI The University of Manchester(曼彻斯特大学) University of Oxford(牛津大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages; Project page at https://alexbubu.com/mobi

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14666 2025-04-22 cs.CV 79%

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Kaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao, Liyu Jia, Wei Zhao, Juncheng Li, Siliang Tang, Hanwang Zhang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(新加坡国立大学) Peking University(北京大学) Huawei Singapore Research Center(华为新加坡研究中心)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04161 2025-04-22 cs.CV 79%

Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model

Keda Tao, Jinjin Gu, Yulun Zhang, Xiucheng Wang, Nan Cheng

机构 * Xidian University(西安电子科技大学) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) State Key Laboratory of Integrated Services Networks(集成服务网络国家重点实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 23 Pages, 28 Figures, ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13631 2025-04-21 cs.AI 79%

Multi-modal Knowledge Graph Generation with Semantics-enriched Prompts

Yajing Xu, Zhiqiang Liu, Jiaoyan Chen, Mingchen Tu, Zhuo Chen, Jeff Z. Pan, Yichi Zhang, Yushan Zhu, Wen Zhang, Huajun Chen

机构 * Zhejiang University(浙江大学) Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) School of Informatics, The University of Edinburgh(爱丁堡大学信息学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08857 2025-04-21 cs.CV 79%

DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation

Minbin Huang, Yanxin Long, Xinchi Deng, Ruihang Chu, Jiangfeng Xiong, Xiaodan Liang, Hong Cheng, Qinglin Lu, Wei Liu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Project page: https://hunyuan-dialoggen.github.io/. Accepted to NAACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12844 2025-04-18 cs.CV 79%

High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion

Libo Zhang, Yongsheng Yu, Jiali Yao, Heng Fan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IJCV. arXiv admin note: text overlap with arXiv:2208.11850

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04110 2025-04-17 cs.HC cs.AI 79%

InterChat: Enhancing Generative Visual Analytics using Multimodal Interactions

Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Piazentin Ono, Wenbin He, Liu Ren, Dongyu Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments This work is accepted by the 27th Eurographics Conference on Visualization (EuroVis 2025). The paper contains 12 pages and 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08358 2025-04-14 cs.CV 79%

LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, Xiongkuo Min

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08111 2025-04-14 cs.CV 79%

POEM: Precise Object-level Editing via MLLM control

Marco Schouten, Mehmet Onurcan Kaya, Serge Belongie, Dim P. Papadopoulos

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to SCIA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13947 2025-04-14 cs.CV 79%

Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation

Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Jiachen Li, Amit K Roy Chowdhury

专题命中 多模态生成 :MLLM(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05821 2025-04-14 cs.CV 79%

F-LMM: Grounding Frozen Large Multimodal Models

Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu, Wentao Liu, Wei Li, Chen Change Loy

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://github.com/wusize/F-LMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03607 2025-04-07 cs.CV 79%

Multimodal Diffusion Bridge with Attention-Based SAR Fusion for Satellite Image Cloud Removal

Yuyang Hu, Suhas Lohit, Ulugbek S. Kamilov, Tim K. Marks

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24165 2025-04-01 cs.LG cs.AI 79%

Predicting Targeted Therapy Resistance in Non-Small Cell Lung Cancer Using Multimodal Machine Learning

Peiying Hua, Andrea Olofson, Faraz Farhadi, Liesbeth Hondelink, Gregory Tsongalis, Konstantin Dragnev, Dagmar Hoegemann Savellano, Arief Suriawinata, Laura Tafe, Saeed Hassanpour

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08650 2025-04-01 cs.CL 79%

An End-to-End Model for Photo-Sharing Multi-modal Dialogue Generation

Peiming Guo, Sinuo Liu, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Min Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20198 2025-03-27 cs.CV 79%

Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models

Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, Min Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16254 2025-03-21 cs.CV 79%

M2N2V2: Multi-Modal Unsupervised and Training-free Interactive Segmentation

Markus Karmann, Peng-Tao Jiang, Bo Li, Onay Urfalioglu

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14216 2025-03-19 cs.LG cs.AI cs.CE 79%

TFG-Flow: Training-free Guidance in Multimodal Generative Flow

Haowei Lin, Shanda Li, Haotian Ye, Yiming Yang, Stefano Ermon, Yitao Liang, Jianzhu Ma

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09901 2025-03-19 cs.CV 79%

MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow

Zhe Li, Yisheng He, Lei Zhong, Weichao Shen, Qi Zuo, Lingteng Qiu, Zilong Dong, Laurence Tianruo Yang, Weihao Yuan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11073 2025-03-17 cs.CV 79%

Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models

Hongyang Wei, Shuaizheng Liu, Chun Yuan, Lei Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10604 2025-03-14 cs.CV 79%

MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction

Yingshuang Zou, Yikang Ding, Chuanrui Zhang, Jiazhe Guo, Bohan Li, Xiaoyang Lyu, Feiyang Tan, Xiaojuan Qi, Haoqian Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09491 2025-03-13 cs.CV eess.IV 79%

DAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution Prediction

Junjie Zhou, Shouju Wang, Yuxia Tang, Qi Zhu, Daoqiang Zhang, Wei Shao

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏