arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2409.01086 2025-09-18 cs.CV cs.AI 81%

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

机构 * Shenzhen Technology University(深圳科技大学) University of Washington(华盛顿大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 81%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08847 2025-09-12 cs.AI cs.CL cs.LG cs.SE 81%

Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs

Amna Hassan

机构 * UET Taxila(塔希尔大学工程学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08489 2025-09-11 cs.CV cs.AI 81%

Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation

Kaleem Ahmad

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07817 2025-09-10 cs.CL cs.MM 81%

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * National University of Singapore(新加坡国立大学) Southern University of Science and Technology(南方科技大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05263 2025-09-09 cs.AI cs.CV cs.LG 81%

LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation

Yinglin Duan, Zhengxia Zou, Tongwei Gu, Wei Jia, Zhan Zhao, Luyi Xu, Xinzhu Liu, Yenan Lin, Hao Jiang, Kang Chen, Shuang Qiu

机构 * NetEase, Inc.(网易公司) Beihang University(北京航空航天大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学) Independent Researcher & Technical Artists(独立研究者及技术艺术家)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05714 2025-09-09 cs.AI cs.CV 81%

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs

Zhaoyu Fan, Kaihang Pan, Mingze Zhou, Bosheng Qin, Juncheng Li, Shengyu Zhang, Wenqiao Zhang, Siliang Tang, Fei Wu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03535 2025-09-05 cs.CL cs.AI 81%

QuesGenie: Intelligent Multimodal Question Generation

Ahmed Mubarak, Amna Ahmed, Amira Nasser, Aya Mohamed, Fares El-Sadek, Mohammed Ahmed, Ahmed Salah, Youssef Sobhy

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 8 figures, 12 tables. Supervised by Dr. Ahmed Salah and TA Youssef Sobhy

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19320 2025-08-29 cs.CV cs.AI 81%

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Ming Chen, Liyuan Cui, Wenyuan Zhang, Haoxian Zhang, Yan Zhou, Xiaohan Li, Songlin Tang, Jiwen Liu, Borui Liao, Hejia Chen, Xiaoqiang Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Technical Report. Project Page: https://chenmingthu.github.io/milm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15194 2025-08-27 cs.CV cs.AI cs.LG 81%

DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Sungnyun Kim, Junsoo Lee, Kibeom Hong, Daesik Kim, Namhyuk Ahn

机构 * KAIST AI(韩国科学技术院人工智能研究所) NAVER WEBTOON AI Sookmyung Women’s University(成均馆女子大学) Inha University(釜山大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Expert Systems with Applications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16930 2025-08-26 eess.AS cs.CV cs.SD 81%

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, Zhao Zhong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03001 2025-08-26 cs.CV cs.MM 81%

One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning

Hao Sun, Yu Song, Jiaqing Liu, Jihong Hu, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Information Science and Engineering, Ritsumeikan University(立命馆大学信息科学与工程学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM 81%

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06492 2025-08-11 cs.CV cs.CL 81%

Effective Training Data Synthesis for Improving MLLM Chart Understanding

Yuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li, Gaowen Liu, Ali Payani, Yuan-Sen Ting, Liang Zheng

机构 * Australian National University(澳大利亚国立大学) Ohio State University(俄亥俄州立大学) Cisco(思科公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025 (poster). 26 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06510 2025-08-08 cs.CV cs.AI 81%

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

Shidan He, Lei Liu, Xiujun Shu, Bo Wang, Yuanhao Feng, Shen Zhao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03069 2025-08-08 cs.CV cs.AI 81%

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K. Du, Zehuan Yuan, Xinglong Wu

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2025; Code and models: https://github.com/ByteVisionLab/TokenFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21167 2025-08-07 cs.CV cs.AI 81%

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * independent researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05683 2025-08-07 cs.LG cs.AI cs.CR cs.MM 81%

Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR

Fardis Nadimi, Payam Abdisarabshali, Kasra Borazjani, Jacob Chakareski, Seyyedali Hosseinalipour

机构 * University at Buffalo–SUNY(布法罗大学-纽约州立大学) Department of Electrical Engineering(电气工程系) New Jersey Institute of Technology (NJIT)(新泽西理工学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 16 pages, 4 Figures, 8 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03426 2025-08-06 cs.CV cs.AI cs.LG 81%

R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation

Futian Wang, Yuhan Qiao, Xiao Wang, Fuling Wang, Yuxiang Zhang, Dengdi Sun

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23058 2025-08-01 cs.CV cs.AI 81%

Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

Alexandru Buburuzan

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments A dissertation submitted to The University of Manchester for the degree of Bachelor of Science in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22920 2025-08-01 cs.CL cs.AI 81%

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey

Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17083 2025-07-24 cs.CV cs.AI 81%

SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction

Zaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An, Junfeng Ding, Jie Zhan, Yunbiao Xu, Jie Ma

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07611 2025-07-15 cs.CL cs.AI 81%

Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Yida Xu, Yunya Song, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海先进研究院) Hong Kong University of Science and Technology(香港科技大学) The University of Manchester(曼彻斯特大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 13 pages. 7 figures

Journal ref This paper is accpeted by ACL2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08719 2025-07-14 cs.CL cs.AI cs.SE 81%

Multilingual Multimodal Software Developer for Code Generation

Linzheng Chai, Jian Yang, Shukai Liu, Wei Zhang, Liran Wang, Ke Jin, Tao Sun, Congnan Liu, Chenchen Zhang, Hualei Zhu, Jiaheng Liu, Xianjie Wu, Ge Zhang, Tianyu Liu, Zhoujun Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20214 2025-07-09 cs.CV cs.MM 81%

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

Yanzhe Chen, Huasong Zhong, Yan Li, Zhenheng Yang

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05626 2025-07-03 cs.CV cs.AI 81%

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

Aarti Ghatkesar, Ganesh Venkatesh

机构 * AppliedML, Cerebras(应用机器学习,Cerebras)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18095 2025-06-24 cs.CV cs.AI cs.LG 81%

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen, Ke Ji, Xidong Wang, Yunjin Yang, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17218 2025-06-23 cs.CV cs.AI 81%

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, Chuang Gan

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Project page: https://vlm-mirage.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12325 2025-06-17 cs.SD cs.CL eess.AS 81%

GSDNet: Revisiting Incomplete Multimodal-Diffusion from Graph Spectrum Perspective for Conversation Emotion Recognition

Yuntao Shou, Jun Yao, Tao Meng, Wei Ai, Cen Chen, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业科技学院) Department of Computer Science, Anhui Normal University(计算机科学系,安徽师范大学) Future Technology Institute, South China University of Technology(未来技术研究院,华南理工大学) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11380 2025-06-16 cs.CV cs.AI 81%

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

Xiaoxin Lu, Ranran Haoran Zhang, Yusen Zhang, Rui Zhang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 18 pages, 10 figures; Accepted to ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏