arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2602.07993 2026-02-10 cs.CV cs.AI 84%

MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance

MCIE: 多模态大语言模型驱动的复杂指令图像编辑与空间引导

Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan, YiFan Zhang, Jack Ma

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 MCIE通过空间引导和背景一致性模块,提升复杂指令图像编辑的指令合规性,实现23.96%的性能提升。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19750 2026-01-29 cs.MM cs.CV cs.IR 84%

Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues

在电子商务目录中评估多模态大语言模型用于缺失模态补全

Junchen Fu, Wenhao Deng, Kaiwen Zheng, Ioannis Arapakis, Yu Ye, Yongxin Ni, Joemon M. Jose, Xuri Ge

机构 * University of Glasgow(格拉斯哥大学) Telefónica Scientific Research(电信科研机构) National University of Singapore(新加坡国立大学) Shandong University(山东大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 本文研究了多模态大语言模型在电子商务目录中补全缺失模态的能力,通过MMPCBench基准测试发现MLLMs在细粒度对齐上存在不足,并探索了GRPO方法以提升补全效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16512 2026-01-29 cs.CV cs.AI 84%

Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection

超越人脸交换:一种基于扩散模型的多模态深度伪造检测基准

Jiaxin Liu, Jia Wang, Saihui Hou, Min Ren, Huijia Wu, Long Ma, Renwang Pei, Zhaofeng He

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DigiFakeAV数据集和DigiShield检测模型,用于多模态深度伪造检测,通过融合时空和跨模态特征提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17761 2026-01-27 cs.LG cs.AI cs.CL 84%

AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation

AR-Omni:一种统一的自回归模型用于任意到任意生成

Dongjie Cheng, Ruifeng Yuan, Yongqi Li, Runyang You, Wenjie Wang, Liqiang Nie, Lei Zhang, Wenjie Li

机构 * The Hong Kong Polytechnic University(香港理工大学) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态生成 :any-to-any(title,abstract);multimodal(abstract);分类 cs.CL、cs.AI

AI总结 AR-Omni提出一种无需专家解码器的统一自回归模型,实现多模态任意到任意生成,解决模态不平衡、视觉保真度和稳定性与创造力平衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15698 2026-01-23 cs.CV cs.AI 84%

Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs

超越视觉安全:通过语义无关输入对多模态大语言模型进行有害图像生成的劫持

Mingyu Yu, Lana Liu, Zhehao Zhao, Wei Wang, Sujuan Qin

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学) School of Cyberspace Security, Beijing University of Posts and Telecommunications(网络安全学院,北京邮电大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文提出BVS框架,通过语义无关输入对多模态大语言模型进行有害图像生成的劫持,揭示其视觉安全边界的脆弱性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03193 2026-01-09 cs.CV cs.AI 84%

UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision

UniCorn:通过自生成监督实现自我改进的统一多模态模型

Ruiyan Han, Zhen Fang, XinYu Sun, Yuchen Ma, Ziheng Wang, Yu Zeng, Zehui Chen, Lin Chen, Wenxuan Huang, Wei-Jie Xu, Yi Cao, Feng Zhao

机构 * MoE Key Lab of BIPC, USTC(脑智能基础与应用联合实验室,中国科学技术大学) FDU(复旦大学) ECNU(华东师范大学) CUHK(香港中文大学) NJU(南京大学) SUDA(上海大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 UniCorn通过自生成监督实现统一多模态模型的自我改进,提升多模态生成质量与理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13107 2025-12-19 cs.CV cs.AI 84%

Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather

基于扩散的多模态3D物体检测在恶劣天气中的修复

Zhijian He, Feifei Liu, Yuwei Li, Zhanpeng Luo, Jintao Cheng, Xieyuanli Chen, Xiaoyu Tang

机构 * School of Xingzhi College, South China Normal University(星智学院,华南师范大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学) School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiffFusion通过基于扩散的修复和自适应跨模态融合,提升多模态3D物体检测在恶劣天气中的鲁棒性与清洁数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15747 2025-12-19 cs.LG cs.CL cs.CV cs.CY 84%

D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models

D3G:多样化的人口数据生成提高多模态模型中的零样本图像分类准确性

Javon Hickmon

机构 * Javon Hickmon(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 D3G通过生成多样化的人口数据,减少多模态模型中的偏见,提高零样本图像分类的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06020 2025-12-09 cs.CV cs.AI 84%

PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation

PrefGen: 多模态偏好学习用于偏好条件下的图像生成

Wenyi Mo, Tianyu Zhang, Yalong Bai, Ligong Han, Ying Ba, Dimitris N. Metaxas

机构 * Rutgers University(罗格斯大学) iN2X MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室) Red Hat AI Innovation(红帽人工智能创新)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 PrefGen通过多模态大语言模型提取用户偏好并注入扩散模型,实现个性化图像生成,优于现有方法。

Comments Project Page: \href{https://prefgen.github.io/}{\texttt{https://prefgen.github.io}}

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22943 2025-12-01 cs.CL cs.CV 84%

Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework

基于成语的视觉双关:一个迭代的LLM-T2IM-MLLM框架

Kelaiti Xiao, Liang Yang, Dongyu Zhang, Paerhati Tulajiang, Hongfei Lin

机构 * Dalian University of Technology(大连理工大学) Xinjiang Normal University(新疆师范大学)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出一个迭代框架,结合LLM、T2IM和MLLM,实现基于成语的视觉双关的自动生成与评估,实验表明MLLM选择对性能影响最大。

Comments Submitted to ICASSP 2026 (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21698 2025-12-01 cs.MM cs.AI 84%

TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement

TIP和Polish:通过共同性-差异性建模与细化的文本-图像-原型引导的多模态生成

Zhiyong Ma, Jiahao Chen, Qingyuan Chuai, Zhengping Li

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

AI总结 TIPPo通过共同性-差异性建模与细化,提升多模态生成的主题一致性和风格一致性。

Comments Submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18780 2025-11-27 cs.CV cs.AI 84%

ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection

ConceptGuard:通过多模态风险检测实现文本-图像到视频生成的主动安全

Ruize Ma, Minghong Cai, Yilei Jiang, Jiaming Han, Yi Feng, Yingshui Tan, Xiaoyong Zhu, Bo Zhang, Bo Zheng, Xiangyu Yue

机构 * CUHK MMLab(香港中文大学多模态实验室) Future Lab, Alibaba Group(阿里巴巴集团未来实验室) Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 ConceptGuard通过多模态风险检测实现文本-图像到视频生成的主动安全,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13243 2025-11-18 cs.LG cs.AI cs.CV 84%

Uncovering and Mitigating Transient Blindness in Multimodal Model Editing

Xiaoqi Han, Ru Li, Ran Yi, Hongye Tan, Zhuomin Liang, Víctor Gutiérrez-Basulto, Jeff Z. Pan

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14245 2025-11-10 cs.CV cs.CL 84%

Towards Explainable Fake Image Detection with Multi-Modal Large Language Models

Yikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen, jun lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, Jianfu Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted to ACM MM 2025; 14 pages including Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24770 2025-11-04 eess.IV cs.AI cs.CV 84%

DMVFC: Deep Learning Based Functionally Consistent Tractography Fiber Clustering Using Multimodal Diffusion MRI and Functional MRI

Bocheng Guo, Jin Wang, Yijie Li, Junyi Wang, Mingyu Gao, Puming Feng, Yuqian Chen, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Lauren J O'Donnell, Fan Zhang

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24820 2025-10-30 cs.CV cs.AI 84%

SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing

Ruiyang Zhang, Jiahao Luo, Xiaoru Feng, Qiufan Pang, Yaodong Yang, Juntao Dai

机构 * PKU Alignment Team, Peking University(北京大学对齐团队) LLM Safety Centre, Beijing Academy of Artificial Intelligence(北京人工智能研究院大语言模型安全中心)

专题命中 多模态生成 :MLLM(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23449 2025-10-28 cs.MM cs.CV cs.IR 84%

CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection

Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou

机构 * School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) National University of Singapore(新加坡国立大学) Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17692 2025-10-17 cs.CL cs.AI cs.LG 84%

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang

机构 * Beihang University(北航) M-A-P The Hong Kong Polytechnic University(香港理工大学) AIWaves University of Alberta(阿尔伯塔大学) University of Waterloo(滑铁卢大学) University of Manchester(曼彻斯特大学) Chinese Academy of Sciences(中国科学院) Peking University(北京大学) Shanghai AI Lab(上海AI实验室) Nanjing University(南京大学) Kuaishou Technology(快手科技)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 (Oral). Codes and models are available in https://github.com/MIO-Team/MIO

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24361 2025-10-01 cs.CV cs.AI cs.HC 84%

UI-UG: A Unified MLLM for UI Understanding and Generation

Hao Yang, Weijie Qiu, Ru Zhang, Zhou Fang, Ruichao Mao, Xiaoyu Lin, Maji Huang, Zhaosong Huang, Teng Guo, Shuoyang Liu, Hai Rao

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10424 2025-09-16 cs.CV cs.AI 84%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08519 2025-09-11 cs.CV cs.MM 84%

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Liyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li, Zhuowei Chen, Lijie Liu, Xu He, Gen Li, Qian He, Zhiyong Wu

机构 * Tsinghua University(清华大学) Intelligent Creation Lab, ByteDance(字节跳动智能创作实验室)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02906 2025-08-21 cs.CL cs.AI 84%

Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement

Zhihan Zhang, Yixin Cao, Lizi Liao

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05234 2025-08-08 cs.CL cs.AI 84%

Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation

Haonan Shangguan, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Ge Yu

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21741 2025-07-30 cs.CV cs.MM 84%

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen

机构 * Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司) Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.MM

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20368 2025-07-29 cs.CV cs.MM 84%

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

Shuolin Xu, Bingyuan Wang, Zeyu Cai, Fangteng Fu, Yue Ma, Tongyi Lee, Hongchuan Yu, Zeyu Wang

机构 * National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学计算机动画国家中心) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与信息工程系)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments 8 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01704 2025-06-04 cs.AI cs.CL 84%

Generate, Not Recommend: Personalized Multimodal Content Generation

Jiongnan Liu, Zhicheng Dou, Ning Hu, Chenyan Xiong

机构 * School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学agate智能学院) Serendipity One Inc.(Serendipity One公司)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14664 2025-05-29 cs.CV cs.AI cs.HC cs.LG 84%

AKRMap: Adaptive Kernel Regression for Trustworthy Visualization of Cross-Modal Embeddings

Yilin Ye, Junchao Huang, Xingchen Zeng, Jiazhi Xia, Wei Zeng

专题命中 多模态生成 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05679 2025-05-01 cs.CV cs.AI 84%

BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents

Yumeng Zhang, Shi Gong, Kaixin Xiong, Xiaoqing Ye, Xiaofan Li, Xiao Tan, Fan Wang, Jizhou Huang, Hua Wu, Haifeng Wang

机构 * Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22941 2025-04-01 cs.AI cs.LG cs.MM 84%

Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

Yugen Sato, Tomohiro Takagi

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10391 2025-03-14 cs.CV cs.AI 84%

CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Yufan Deng, Xun Guo, Yizhi Wang, Jacob Zhiyuan Fang, Angtian Wang, Shenghai Yuan, Yiding Yang, Bo Liu, Haibin Huang, Chongyang Ma

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏