arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2510.08530 2025-10-10 cs.GR cs.CV 79%

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang, Hao Zhu, Jing Liao

机构 * City University of Hong Kong(香港城市大学) WeChat, Tencent Inc(微信、腾讯公司) Manycore Tech Inc(很多核科技公司)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06679 2025-10-09 cs.CV 79%

DreamOmni2: Multimodal Instruction-based Editing and Generation

Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang, Jiyang Liu, Jingyao Li, Haoru Tan, Sitong Wu, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港科技大学) HKU(香港大学) ByteDance Inc(字节跳动公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06308 2025-10-09 cs.CV 79%

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Yi Xin, Qi Qin, Siqi Luo, Kaiwen Zhu, Juncheng Yan, Yan Tai, Jiayi Lei, Yuewen Cao, Keqi Wang, Yibin Wang, Jinbin Bai, Qian Yu, Dengyang Jiang, Yuandong Pu, Haoxing Chen, Le Zhuo, Junjun He, Gen Luo, Tianbin Li, Ming Hu, Jin Ye, Shenglong Ye, Bo Zhang, Chang Xu, Wenhai Wang, Hongsheng Li, Guangtao Zhai, Tianfan Xue, Bin Fu, Xiaohong Liu, Yu Qiao, Yihao Liu

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Nanjing University(南京大学) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 33 pages, 13 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04765 2025-10-07 cs.AI 79%

LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0

Jinbo Wen, Jiawen Kang, Linfeng Zhang, Xiaoying Tang, Jianhang Tang, Yang Zhang, Zhaohui Yang, Dusit Niyato

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Guangdong University of Technology(广东工业大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) Guizhou University(贵州大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02880 2025-10-06 cs.AI 79%

Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

Tianren Ma, Mu Zhang, Yibing Wang, Qixiang Ye

机构 * University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Project Page: https://github.com/martian422/MaskGRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00015 2025-10-03 cs.CL 79%

Design and Application of Multimodal Large Language Model Based System for End to End Automation of Accident Dataset Generation

MD Thamed Bin Zaman Chowdhury, Moazzem Hossain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00647 2025-10-02 cs.CL 79%

MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation

Jinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang, Xiaoyu Shen, Xipeng Qiu, See-Kiong Ng

机构 * National University of Singapore(新加坡国立大学) Fudan University(复旦大学) Harbin Institute of Technology(哈尔滨工业大学) Eastern Institute of Technology(东方技术研究所)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08678 2025-10-02 cs.CV 79%

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

Juan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin, Taesup Kim

专题命中 多模态生成 :any-to-any(title,abstract);分类 cs.CV

Comments Accepted at ICCV25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05689 2025-10-02 cs.CV 79%

GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

Zebin Xing, Xingyu Zhang, Yang Hu, Bo Jiang, Tong He, Qian Zhang, Xiaoxiao Long, Wei Yin

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Horizon Robotics Nanjing University(南京大学) Huazhong University of Science & Technology(华中科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26641 2025-10-01 cs.CV 79%

Query-Kontext: An Unified Multimodal Model for Image Generation and Editing

Yuxin Song, Wenkai Dong, Shizun Wang, Qi Zhang, Song Xue, Tao Yuan, Hu Yang, Haocheng Feng, Hang Zhou, Xinyan Xiao, Jingdong Wang

机构 * Baidu VIS(百度视觉) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22930 2025-09-30 cs.CV 79%

FishAI 2.0: Marine Fish Image Classification with Multi-modal Few-shot Learning

Chenghan Yang, Peng Zhou, Dong-Sheng Zhang, Yueyun Wang, Hong-Bin Shen, Xiaoyong Pan

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22570 2025-09-29 cs.AI 79%

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

Qi Mao, Tinghan Yang, Jiahao Li, Bin Li, Libiao Jin, Yan Lu

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室) Communication University of China(中国传媒大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18824 2025-09-24 cs.CV 79%

Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation

Yanzuo Lu, Xin Xia, Manlin Zhang, Huafeng Kuang, Jianbin Zheng, Yuxi Ren, Xuefeng Xiao

机构 * ByteDance Seed(字节跳动种子)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15246 2025-09-22 cs.GR cs.AI 79%

GenCAD-3D: CAD Program Generation using Multimodal Latent Space Alignment and Synthetic Dataset Balancing

Nomi Yu, Md Ferdous Alam, A. John Hart, Faez Ahmed

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 9 figures, 15 pages. Accepted and soon published in the ASME Journal of Mechanical Design

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09315 2025-09-17 cs.RO cs.CV cs.LG 79%

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving

Xuefeng Jiang, Yuan Ma, Pengxiang Li, Leimeng Xu, Xin Wen, Kun Zhan, Zhongpu Xia, Peng Jia, Xianpeng Lang, Sheng Sun

机构 * LiAuto Inc(LiAuto公司) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动研究所)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11082 2025-09-16 cs.CV cs.RO 79%

Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation

Zongwu Xie, Kaijie Yun, Yang Liu, Yiming Ji, Han Li

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07473 2025-09-10 cs.AI 79%

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

Qin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu, Han Shi, Dongmei Zhang

机构 * Peking University(北京大学) Microsoft(微软)

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12528 2025-09-09 cs.CV 79%

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19065 2025-09-08 cs.CV 79%

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

Zhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学科学与技术大学) Lanzhou University(兰州大学) Meta AI The University of Sydney(悉尼大学) IHPC, A*STAR(IHPC,A*STAR)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments ICCV 2025, Project in https://wikiautogen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01074 2025-09-03 cs.CV cs.GR 79%

Multimodal Conditional 3D Face Geometry Generation

Christopher Otto, Prashanth Chandran, Sebastian Weiss, Markus Gross, Gaspard Zoss, Derek Bradley

机构 * ETH Zürich(苏黎世联邦理工学院) DisneyResearch | Studios(迪士尼研究室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Added more evaluation since the first version. Accepted to SMI 2025. Computers & Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13602 2025-09-03 cs.CV 79%

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

Xiaolu Hou, Bing Ma, Jiaxiang Cheng, Xuhua Ren, Kai Yu, Wenyue Li, Tianxiang Zheng, Qinglin Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://personavlog-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21460 2025-09-01 cs.IR cs.AI 79%

Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate Prediction

Xiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li, Zhejun Zhao

机构 * Peking University(北京大学) Wuhan University(武汉大学) Shanghai University of International Business(上海国际商务大学) Microsoft(微软)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments SIGIR 2025

Journal ref SIGIR 2025: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval Pages 581 - 591

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20379 2025-08-29 cs.CV 79%

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

Hyeonyu Kim, Seokhoon Jeong, Seonghee Han, Chanhyuk Choi, Taehwan Kim

机构 * MAUM AI Inc.(MAUM AI公司) Artificial Intelligence Graduate School UNIST(UNIST人工智能研究生学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09242 2025-08-28 cs.AI 79%

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, Soroosh Tayebi Arasteh

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref Biomed. Eng. Lett. 15 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17614 2025-08-26 cs.CV 79%

JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

Aowen Wang, Wei Li, Hao Luo, Mengxing Ao, Chenyu Zhu, Xinyang Li, Fan Wang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(汇安实验室) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17199 2025-08-26 cs.CV 79%

MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling

Hyeyeon Kim, Sungwoo Han, Jingun Kwon, Hidetaka Kamigaito, Manabu Okumura

机构 * Chungnam National University(Chungnam 国立大学) Nara Institute of Science and Technology (NAIST)(Nara 科学技术研究所) Institute of Science Tokyo(东京科学研究所)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16763 2025-08-26 cs.CV 79%

WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar

机构 * ServiceNow Mila Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted to the EMNLP 2025 main conference. Check the project page here: https://webmmu-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13068 2025-08-19 cs.CV cs.LG 79%

Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy, Farig Sadeque, Swakkhar Shatabda

机构 * Department of Computer Science and Engineering, Brac University(计算机科学与工程系,布拉克大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12399 2025-08-19 cs.CV 79%

Federated Cross-Modal Style-Aware Prompt Generation

Suraj Prasad, Navyansh Mahla, Sunny Gupta, Amit Sethi

机构 * Indian Institute of Technology Bombay(印度理工学院孟买学院)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10118 2025-08-19 cs.LG cs.CV 79%

From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation

Ke Niu, Haiyang Yu, Zhuofan Chen, Mengyang Zhao, Teng Fu, Bin Li, Xiangyang Xue

机构 * Fudan University(复旦大学) ByteDance Inc(字节跳动公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏