arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2603.08021 2026-03-31 cs.RO cs.CV 74%

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

AffordGrasp:跨模态扩散用于感知意识抓取合成

Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yujiao Shi, Xuming He

机构 * ShanghaiTech University(上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

AI总结 本文提出AffordGrasp,通过跨模态扩散模型生成准确反映物体几何和用户指令的抓取姿态,提升AR/VR和具身AI中的手-物交互质量。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20060 2026-03-27 cs.CV cs.RO 74%

MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving

MeanFuser: 一种基于MeanFlow的高效多模态轨迹生成与自适应重构方法用于端到端自动驾驶

Junli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing, Pengfei Li, Guang Li, Kun Ma, Guang Chen, Hangjun Ye, Zhongpu Xia, Long Chen, Qichao Zhang

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Xiaomi EV(小米汽车) Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

AI总结 本文提出MeanFuser,通过引入Gaussian Mixture Noise、MeanFlow Identity和轻量级ARM模块,实现了高效且鲁棒的多模态轨迹生成与自适应重构,提升了端到端自动驾驶的性能和效率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08998 2026-03-11 cs.CV 74%

Diffusion-Based Authentication of Copy Detection Patterns: A Multimodal Framework with Printer Signature Conditioning

基于扩散的复制检测模式认证:一种多模态框架与打印机签名条件化

Bolutife Atoki, Iuliia Tkachenko, Bertrand Kerautret, Carlos Crispim-Junior

机构 * Université Lumière Lyon 2, CNRS, INSA Lyon, Universite Claude Bernard Lyon 1, LIRIS UMR5205(里摩日大学里昂2分校、法国国家科学研究中心、里昂国立应用科学学院、里昂大学克莱尔-贝尔纳分校、LIRIS UMR5205)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 本文提出基于扩散的多模态认证框架,利用打印机签名条件化技术,实现对复制检测模式的高效认证,优于传统方法和深度学习方法。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18903 2026-02-24 cs.CV cs.HC 74%

SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model

Gemini 3 Pro图像的SCHEMA:为Google原生多模态模型的受控AI图像生成的结构化方法

Luca Cazzaniga

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 SCHEMA为Google Gemini 3 Pro图像提供结构化提示工程方法,通过三级系统提升AI图像生成的可控性,实现高合规率和跨领域应用。

Comments 24 pages, 8 tables. Based on SCHEMA Method v1.0 (deposited December 11, 2025). Previously published on Zenodo: doi:10.5281/zenodo.18721380

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23208 2026-02-03 cs.CL 74%

A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks

一种评估和增强多模态大语言模型在文化情境任务中解释能力的结构框架

Haorui Yu, Ramon Ruiz-Dolz, Qiufeng Yi

机构 * DJCAD, University of Dundee, United Kingdom(邓迪大学DJCAD部门) ARG-tech, SSEN, University of Dundee, United Kingdom(邓迪大学) School of Computer Science, University of Birmingham, United Kingdom(伯明翰大学计算机科学学院)

专题命中 多模态生成 :multimodal(title);分类 cs.CL

AI总结 本研究提出了一种结构框架,用于评估和增强多模态大语言模型在文化情境任务中生成中国绘画批评的能力,通过量化评价特征和人设引导提示,揭示了VLMs在艺术批评领域的表现与局限。

Comments EMNLP 2025 submission, 10 pages, 6 figures, 5 tables

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 1945-1971, Suzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19774 2026-01-22 cs.LG cs.AI eess.SP 74%

PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection

PPGFlowECG: 基于跨模态编码的潜在修正流用于PPG引导的ECG生成和心血管疾病检测

Xiaocheng Fang, Jiarui Jin, Haoyu Wang, Che Liu, Jieyi Cai, Yujie Xiao, Guangkun Nie, Bo Liu, Shun Huang, Hongyan Li, Shenda Hong

机构 * National Institute of Health Data Science, Peking University, China(北京大学国家健康数据科学研究院) School of Intelligence Science and Technology, Peking University, China(北京大学智能科学与技术学院) Data Science Institute, Imperial College London, UK(伦敦帝国理工学院数据科学研究院) University of Chinese Academy of Sciences, China(中国科学院大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.AI

AI总结 PPGFlowECG通过跨模态编码和潜在修正流,实现PPG引导的ECG生成,提升心血管疾病检测的可扩展性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04506 2026-01-09 cs.LG cs.AI cs.CE 74%

Surface-based Molecular Design with Multi-modal Flow Matching

基于表面的分子设计与多模态流匹配

Fang Wu, Zhengyuan Zhou, Shuting Jin, Xiangxiang Zeng, Jure Leskovec, Jinbo Xu

机构 * Stanford University(斯坦福大学) University of California, San Diego(加州大学圣地亚哥分校) Wuhan University of Science and Technology(武汉科技大学) Hunan University(湖南大学)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

AI总结 SurfFlow通过多模态流匹配算法实现基于分子表面的肽共同设计,提升肽与受体的结合准确性,并在PepMerge基准中优于全原子基线。

Journal ref KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00477 2025-12-24 cs.CV 74%

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

LAMIC:基于多模态扩散变换器可扩展性的布局感知多图像合成

Yuzhuo Chen, Zehua Ma, Jianhua Wang, Kai Kang, Shunyu Yao, Weiming Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Onestory Team(Onestory团队) East China Normal University(华东师范大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 LAMIC通过无训练方式实现多参考图像合成,提升布局控制与背景一致性。

Comments 8 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07934 2025-11-12 cs.CV 74%

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo, Hongyuan Zhang

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07816 2025-11-12 cs.CV 74%

Cancer-Net PCa-MultiSeg: Multimodal Enhancement of Prostate Cancer Lesion Segmentation Using Synthetic Correlated Diffusion Imaging

Jarett Dewbury, Chi-en Amy Tai, Alexander Wong

机构 * Systems Design Engineering University of Waterloo(水力工程系统设计系大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted at ML4H 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22439 2025-10-30 cs.SD cs.AI 74%

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Ali Vosoughi, Yongyi Zang, Qihui Yang, Nathan Paek, Randal Leistikow, Chenliang Xu

机构 * Smule Labs(Smule实验室) University of California, San Diego(加州大学圣地亚哥分校) University of Rochester(罗切斯特大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments 9 pages, 2 figures, 4 tables; v2: corrected spelling of a co-author name; no content changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22431 2025-10-28 cs.MA cs.CV 74%

Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration

Zheng Wei, Mingchen Li, Zeqian Zhang, Ruibin Yuan, Pan Hui, Huamin Qu, James Evans, Maneesh Agrawala, Anyi Rao

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Chicago(芝加哥大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23635 2025-09-30 cs.CV 74%

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16768 2025-09-23 cs.CV 74%

MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation

Omid Bonakdar, Nasser Mozayani

机构 * School of Computer engineering, Iran university of Science and Technology(计算机工程学院,伊朗科学技术大学)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13227 2025-09-18 math.OC cs.AI cs.SY eess.SY 74%

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Santanu Banerjee, Goutam Sen, Siddhartha Mukhopadhyay

机构 * Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

Comments Major changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20877 2025-09-12 cs.CV 74%

Deep Learning Framework for Early Detection of Pancreatic Cancer Using Multi-Modal Medical Imaging Analysis

Dennis Slobodzian, Amir Kordijazi

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

Comments 21 pages, 17 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16376 2025-09-09 cs.CV 74%

LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images

Leyang Wang, Joice Lin

机构 * University College London(伦敦大学学院) Xiamen University(厦门大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08987 2025-08-13 cs.CV cs.HC 74%

ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation

Ding Xia, Naoto Inoue, Qianru Qiu, Kotaro Kikuchi

机构 * The University of Tokyo(东京大学) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted to ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23676 2025-08-01 cs.LG cs.CV 74%

DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data

Rabeya Tus Sadia, Qiang Cheng

机构 * Department of Computer Science University of Kentucky(计算机科学系 哥伦比亚大学) Department of Computer Science, Institute for Biomedical Informatics University of Kentucky(计算机科学系 生物医学信息学研究所 哥伦比亚大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21260 2025-07-30 cs.LG cs.AI q-bio.QM 74%

Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors

Amartya Banerjee, Xingyu Xu, Caroline Moosmüller, Harlin Lee

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments Code: https://github.com/amartya21/Adam-PnP

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11571 2025-07-21 cs.CV 74%

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

Jiachen Yu, Yufei Zhan, Ziheng Wu, Yousong Zhu, Jinqiao Wang, Minghui Qiu

机构 * Tsinghua University(清华大学) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ByteDance China(字节跳动中国) Peng Cheng Laboratory(鹏城实验室) Wuhan AI Research(武汉人工智能研究)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05568 2025-07-09 cs.CV cs.LG 74%

ReLayout: Integrating Relation Reasoning for Content-aware Layout Generation with Multi-modal Large Language Models

Jiaxu Tian, Xuehui Yu, Yaoxing Wang, Pan Wang, Guangqian Guo, Shan Gao

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11182 2025-06-16 q-bio.GN cs.AI 74%

Multimodal Modeling of CRISPR-Cas12 Activity Using Foundation Models and Chromatin Accessibility Data

Azim Dehghani Amirabad, Yanfei Zhang, Artem Moskalev, Sowmya Rajesh, Tommaso Mansi, Shuwei Li, Mangal Prakash, Rui Liao

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments This manuscript has been accepted by ICML workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08468 2025-06-10 cs.RO cs.CV 74%

Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation

Haosheng Li, Weixin Mao, Weipeng Deng, Chenyu Meng, Haoqiang Fan, Tiancai Wang, Yoshie Osamu, Ping Tan, Hongan Wang, Xiaoming Deng

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Waseda University(早稻田大学) University of Hong Kong(香港大学) MEGVII Technology(美格智能科技) Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10495 2025-05-26 cs.LG cs.CL 74%

RouteNator: A Router-Based Multi-Modal Architecture for Generating Synthetic Training Data for Function Calling LLMs

Vibha Belavadi, Tushar Vatsa, Dewang Sultania, Suhas Suresha, Ishita Verma, Cheng Chen, Tracy Holloway King, Michael Friedrich

机构 * Adobe Inc.(Adobe公司)

专题命中 多模态生成 :multi-modal(title);分类 cs.CL

Comments Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing

Journal ref https://aclanthology.org/2025.knowledgenlp-1.10/ KnowledgeNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01487 2025-05-07 cs.AI 74%

FastRM: An efficient and automatic explainability framework for multimodal generative models

Gabriela Ben-Melech Stan, Estelle Aflalo, Man Luo, Shachar Rosenman, Tiep Le, Sayak Paul, Shao-Yen Tseng, Vasudev Lal

机构 * Intel Labs(英特尔实验室) Hugging Face

专题命中 多模态生成 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02542 2025-04-08 cs.CV 74%

Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Fa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou, Xiu Li, Qin Lin, Qinglin Lu, Dan Xu

专题命中 多模态生成 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10639 2025-03-14 cs.CV 74%

GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Rongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang, Hao Li, Shilin Yan, Hao Tian, Xingyu Zeng, Rui Zhao, Jifeng Dai, Xihui Liu, Hongsheng Li

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Dataset and models are released in https://github.com/rongyaofang/GoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08165 2025-03-12 cs.CV 74%

Multimodal Generation of Animatable 3D Human Models with AvatarForge

Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang

专题命中 多模态生成 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00619 2025-03-04 cs.IR cs.AI cs.LG 74%

PinLanding: Content-First Keyword Landing Page Generation via Multi-Modal AI for Web-Scale Discovery

Faye Zhang, Jasmine Wan, Qianyu Cheng, Jinfeng Rao

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏