arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2602.19409 2026-02-24 cs.SD 82%

AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration

AuditoryHuM: 利用人机协同生成和聚类听觉场景标签

Henry Zhong, Jörg M. Buchholz, Julian Maclaren, Simon Carlile, Richard F. Lyon

机构 * Australian Hearing Hub, Macquarie University, Sydney, Australia(澳大利亚听力中心、麦觉里大学、悉尼、澳大利亚) Google Research Australia, Sydney, Australia(谷歌澳大利亚研究、悉尼、澳大利亚)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract)

AI总结 AuditoryHuM通过人机协同方法实现听觉场景标签的自动生成与聚类,提供了一种高效且低成本的标准化分类解决方案,适用于边缘设备部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02590 2026-02-04 cs.RO 82%

StepNav: Structured Trajectory Priors for Efficient and Multimodal Visual Navigation

StepNav: 为高效且多模态视觉导航引入结构化轨迹先验

Xubo Luo, Aodi Wu, Haodong Han, Xue Wan, Wei Zhang, Leizheng Shu, Ruisuo Wang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract)

AI总结 StepNav通过结构化多模态轨迹先验提升视觉导航的鲁棒性、效率和安全性。

Comments 8 pages, 7 figures; Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06792 2026-01-13 cs.LG 82%

Cross-Modal Computational Model of Brain-Heart Interactions via HRV and EEG Feature

通过HRV和EEG特征实现脑心交互的跨模态计算模型

Malavika Pradeep, Akshay Sasi, Nusaibah Farrukh, Rahul Venugopal, Elizabeth Sherly

机构 * Digital University Kerala(德里大学凯拉尔) Centre for Consciousness Studies, NIMHANS(意识研究学院,NIMHANS)

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract)

AI总结 本研究通过ECG和EEG特征构建跨模态模型,探索ECG信号作为认知负荷替代指标的可能性,并利用合成数据提升模型鲁棒性。

Comments 6 pages, 2 figures, Code available at: https://github.com/Malavika-pradeep/Computational-model-for-Brain-heart-Interaction-Analysis. Presented at AIHC (not published)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24513 2026-01-05 cs.CY 82%

From Static to Dynamic: Evaluating the Perceptual Impact of Dynamic Elements in Urban Scenes via MLLM-Guided Generative Inpainting

从静态到动态:通过MLLM引导的生成修复技术评估城市场景中动态元素的感知影响

Zhiwei Wei, Mengzi Zhang, Boyan Lu, Zhitao Deng, Nai Yang, Hua Liao

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract)

AI总结 通过MLLM引导的生成修复技术,研究评估了动态元素在城市场景中的感知影响,发现移除动态元素导致活力显著下降,揭示了光照、人类存在和深度变化对感知变化的关键作用。

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08257 2025-12-10 cs.LG eess.IV 82%

Geometric-Stochastic Multimodal Deep Learning for Predictive Modeling of SUDEP and Stroke Vulnerability

几何-随机多模态深度学习用于SUDEP和中风易感性的预测建模

Preksha Girish, Rachana Mysore, Mahanthesha U, Shrey Kumar, Misbah Fatimah Annigeri, Tanish Jain

机构 * Dept. of Artificial Intelligence & Machine Learning(人工智能与机器学习系) B.N.M Institute of Technology(B.N.M技术学院) Dept. of Computer Science Engineering(计算机科学与工程系)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出几何-随机多模态深度学习框架,通过整合多种生物信号建模SUDEP和中风易感性,提升预测准确性和可解释性生物标志物。

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07184 2025-12-09 cs.LG 82%

UniDiff: A Unified Diffusion Framework for Multimodal Time Series Forecasting

UniDiff: 一种用于多模态时间序列预测的统一扩散框架

Da Zhang, Bingyu Li, Zhuyuan Zhao, Junyu Gao, Feiping Nie, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University, Xi’an 710072, China(人工智能学院、光学与电子学(iOPEN)、西北工业大学,西安710072,中国) Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究所(TeleAI)、中国电信,中国)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 UniDiff提出了一种统一的扩散框架,通过融合文本和时间戳信息,提升多模态时间序列预测的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19051 2025-12-05 cs.CV cs.AI cs.MM 82%

Multimodal Markup Document Models for Graphic Design Completion

多模态标记文档模型用于图形设计完成

Kotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani, Edgar Simo-Serra, Kota Yamaguchi

机构 * Waseda University(早稻田大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 MarkupDM是一种多模态标记文档模型,通过统一处理图形设计任务,实现文本、图像和属性值的完成,展示了其在设计自动化中的广泛适用性。

Comments Accepted by ACM Multimedia 2025, Project page: https://cyberagentailab.github.io/MarkupDM/

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia. 2025. p.11022-11031

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23321 2025-12-01 cs.SE 82%

Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing

Chart2Code-MoLA: 通过自适应专家路由实现高效的多模态代码生成

Yifei Wang, Jacky Keung, Zhenyu Mao, Jingyu Zhang, Yuchen Cao

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

AI总结 Chart2Code-MoLA通过自适应专家路由实现高效多模态代码生成,提升生成准确性、降低内存消耗并加速收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17534 2025-11-21 cs.CV cs.CL cs.MM 82%

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

协同强化学习用于统一多模态理解和生成

Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本文提出CoRL框架,通过协同强化学习提升多模态大语言模型在生成与理解任务上的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06284 2025-11-11 cs.CV cs.CL cs.MM 82%

Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

Bing Wang, Ximing Li, Yanjun Wang, Changchun Li, Lin Yuanbo Wu, Buyu Wang, Shengsheng Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by AAAI 2026. 13 pages, 6 figures. Code: https://github.com/wangbing1416/RETSIMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04012 2025-11-07 cs.SE 82%

PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models

Yongxi Chen, Lei Chen

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17002 2025-10-21 cs.LG 82%

EEschematic: Multimodal-LLM Based AI Agent for Schematic Generation of Analog Circuit

Chang Liu, Danial Chitnis

机构 * School of Engineering The University of Edinburgh Edinburgh, UK(工程学院 苏格兰爱丁堡大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16990 2025-10-21 cs.LG 82%

Graph4MM: Weaving Multimodal Learning with Structural Information

Xuying Ning, Dongqi Fu, Tianxin Wei, Wujiang Xu, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Meta AI Rutgers University(罗格斯大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13804 2025-10-16 cs.CV cs.AI cs.CL 82%

Generative Universal Verifier as Multimodal Meta-Reasoner

Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao, Renrui Zhang, Ruihang Chu, Ling Yang, Yujiu Yang

机构 * Tsinghua University(清华大学) ByteDance Seed(字节跳动种子) Princeton University(普林斯顿大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12254 2025-10-15 cs.LG 82%

FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning

Ningxin He, Yang Liu, Wei Sun, Xiaozhou Ye, Ye Ouyang, Tiegang Gao, Zehui Zhang

机构 * Institute for AI Industry Research, Tsinghua University(人工智能产业研究院,清华大学) School of Software Engineering, Nankai University(软件工程学院,南开大学) Department of Computing, Hong Kong Polytechnic University(计算机学院,香港理工大学) AsiaInfo Technologies(亚信息科技) China-Austria Belt and Road Joint Laboratory on Artificial Intelligence and Advanced Manufacturing, Hangzhou Dianzi University(人工智能与先进制造联合实验室,杭州电子科技大学)

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10037 2025-10-14 cs.CE 82%

Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration

Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee, Karanjit Kooner, Jui-Ka Wang, Ning Zhang, Jia Zhang

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by IEEE 25th BIBE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20629 2025-10-06 cs.CV cs.AI cs.MM 82%

AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Pohang University of Science and Technology(釜山科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06561 2025-09-24 cs.CL cs.AI cs.CV 82%

LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles

Ho Yin 'Sam' Ng, Ting-Yao Hsu, Aashish Anantha Ramakrishnan, Branislav Kveton, Nedim Lipka, Franck Dernoncourt, Dongwon Lee, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ting-Hao 'Kenneth' Huang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings. The LaMP-CAP dataset is publicly available at: https://github.com/Crowd-AI-Lab/lamp-cap

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07963 2025-09-09 cs.AI cs.CL cs.CV 82%

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu, Ji-Rong Wen, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) School of Computer Science(计算机科学学院) Baidu Inc.(百度公司) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09945 2025-08-14 cs.CL cs.AI cs.CV 82%

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

Lingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li, Dongdong Zhang, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16326 2025-08-05 cs.LG 82%

ChemMLLM: Chemical Multimodal Large Language Model

Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01735 2025-07-03 cs.CV cs.AI cs.CL cs.LG 82%

ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

Kai Chen, Ruiyuan Gao, Lanqing Hong, Hang Xu, Xu Jia, Holger Caesar, Dengxin Dai, Bingbing Liu, Dzmitry Tsishkou, Songcen Xu, Chunjing Xu, Qiang Xu, Huchuan Lu, Dit-Yan Yeung

机构 * Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学) Dalian University of Technology(大连理工大学) TU Delft(代尔夫特理工大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huawei Zurich Research Center(华为苏黎世研究中心) Huawei IAS BU(华为IAS业务单元)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ECCV 2024. Workshop page: https://coda-dataset.github.io/w-coda2024/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02242 2025-06-19 cs.LG cs.CY 82%

From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models

Yihong Tang, Ao Qu, Xujing Yu, Weipeng Deng, Jun Ma, Jinhua Zhao, Lijun Sun

机构 * McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) The University of Hong Kong(香港大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10007 2025-06-13 cs.MM cs.AI cs.CV 82%

Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space

Kangwei Liu, Junwu Liu, Xiaowei Yi, Jinlin Guo, Yun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Laboratory for Big Data and Decision, School of Systems Engineering, National University of Defense Technology(国防科技大学系统工程学院大数据与决策实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08480 2025-06-11 cs.CL cs.AI cs.CV 82%

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究院,北京大学)

专题命中 多模态生成 :image-text(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22000 2025-05-29 eess.IV 82%

Collaborative Learning for Unsupervised Multimodal Remote Sensing Image Registration: Integrating Self-Supervision and MIM-Guided Diffusion-Based Image Translation

Xiaochen Wei, Weiwei Guo, Wenxian Yu

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19874 2025-05-27 cs.CV cs.AI cs.MM 82%

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

Yi Wu, Lingting Zhu, Shengju Qian, Lei Liu, Wandi Qiao, Lequan Yu, Bin Li

机构 * University of Science and Technology of China(中国科学技术大学) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17613 2025-05-26 cs.AI cs.CL cs.CV 82%

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

Jihan Yao, Yushi Hu, Yujie Yi, Bin Han, Shangbin Feng, Guang Yang, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14011 2025-04-22 cs.CV cs.AI cs.MM 82%

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

Fulvio Sanguigni, Davide Morelli, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08641 2025-04-14 cs.CV cs.AI cs.CL 82%

Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization

Jialu Li, Shoubin Yu, Han Lin, Jaemin Cho, Jaehong Yoon, Mohit Bansal

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Website: https://video-msg.github.io; The first three authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏