arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

1806.09882 2021-03-11 cs.CV 83%

Multi-modal Image Processing based on Coupled Dictionary Learning

Pingfan Song, Miguel R. D. Rodrigues

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments SPAWC 2018, 19th IEEE International Workshop On Signal Processing Advances In Wireless Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04830 2021-02-10 cs.CL 83%

Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment Analysis

Wenmeng Yu, Hua Xu, Ziqi Yuan, Jiele Wu

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments Accepted by AAAI2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.05597 2021-01-19 eess.IV cs.CV cs.LG 83%

EMIXER: End-to-end Multimodal X-ray Generation via Self-supervision

Siddharth Biswal, Peiye Zhuang, Ayis Pyrros, Nasir Siddiqui, Sanmi Koyejo, Jimeng Sun

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00663 2020-05-05 cs.CL 83%

Benchmarking Multimodal Regex Synthesis with Complex Structures

Xi Ye, Qiaochu Chen, Isil Dillig, Greg Durrett

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments ACL2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03548 2019-07-09 cs.CV eess.IV 83%

Unified Attentional Generative Adversarial Network for Brain Tumor Segmentation From Multimodal Unpaired Images

Wenguang Yuan, Jia Wei, Jiabing Wang, Qianli Ma, Tolga Tasdizen

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV

Comments 9 pages, 4 figures, Accepted by MICCAI2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.09915 2018-04-27 cs.CV 83%

Boosting LiDAR-based Semantic Labeling by Cross-Modal Training Data Generation

Florian Piewak, Peter Pinggera, Manuel Schäfer, David Peter, Beate Schwarz, Nick Schneider, David Pfeiffer, Markus Enzweiler, Marius Zöllner

专题命中 多模态生成 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12876 2026-08-14 cs.CV cs.AI 新提交 82%

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

SPARED:基于推理的AI生成图像检测方法,采用对抗编辑数据

Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan

专题命中 多模态生成 :MLLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出对抗强化学习框架SPARED,通过扩散图像编辑器与推理型MLLM的交替博弈,训练出能抗捷径、泛化能力强的AI生成图像检测器,在三个外部基准上性能单调提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11670 2026-06-11 cs.CV cs.AI 新提交 82%

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

ARGUS: 堆叠多视角身份马赛克注入用于主体保持的视频生成

Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan

机构 * Peking University(北京大学) Kuaishou Technology(快手科技) Xiamen University(厦门大学)

专题命中 多模态生成 :MLLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 提出ARGUS框架,通过堆叠多视角身份马赛克注入(SMII)将身份表示为紧凑动态分布,结合MLLM身份导演、无交叉对反事实训练等模块,在主体保持视频生成中达到SOTA。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13679 2026-08-17 cs.GR 新提交 82%

Strand-based Hairstyle Generation via Large Reconstruction and Multimodal Models

基于 strand 的发型生成:结合大型重建模型与多模态模型

Conghui Hao, Tao Huang, Yuefan Shen, Tongtong Wang, Zhongtian Zheng, Kui Wu

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 该研究提出一种结合 LRMs、LMMs 与经典几何处理的自动流程,可从单视图图像快速生成高质量、适配生产的各类基于 strand 的发型,适用于现代数字人工作流。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03922 2026-08-12 cs.CL cs.AI cs.CV 版本更新 82%

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

HSSBench: 多模态大语言模型在人文学与社会科学能力评估中的基准测试

Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, Xuelong Li

机构 * National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心) Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院) Tsinghua University(清华大学) Chinese Academy of Sciences(中国科学院) University of British Columbia(不列颠哥伦比亚大学) Renmin University of China(中国人民大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 HSSBench是一个专门评估多模态大语言模型在人文学与社会科学任务能力的基准测试,包含多语言样本,通过协作生成数据提升跨学科推理能力。

Comments ICLR 2026 (OpenReview: https://openreview.net/forum?id=iQsKotob31)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00410 2026-08-04 cs.AI cs.CL cs.CV 新提交 82%

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

歧义去了哪里?探究多模态模型如何解释多义词

Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths

机构 * Princeton University(普林斯顿大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。

Comments Oral Presentation, Sci-FM Workshop @ COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08497 2026-07-10 cs.CV cs.AI cs.CL cs.LG 新提交 82%

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

用于多模态理解、生成和编辑的认知结构多模态智能体

Feng Wang, Canmiao Fu, Zhipeng Huang, Chen Li, Jing Lyu, Ge Li

机构 * Peking University(北京大学) Tencent Inc(腾讯公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对统一多模态模型在长时多模态对话中的局限,提出认知结构多模态智能体,通过外化视觉信息等方式改进。开发统一场景引擎,构建评估基准。实验显示该智能体检索准确率高、推理时间减半,还介绍了工具包,为长时多模态智能体提供新范式。

Comments 16 pages, 7 figures, 8 tables. Project page: https://caseclose.github.io/cma-harness/ Code: https://github.com/caseclose/cma-harness

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22469 2026-07-09 eess.SP 版本更新 82%

Multi-Modal Beamforming with Model Compression and Modality Generation for V2X Networks

用于车联网的基于模型压缩和模态生成的多模态波束成形

Chen Shang, Dinh Thai Hoang, Jiadong Yu

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 针对6G车联网中ISAC范式的不足,提出基于分层Transformer的多模态学习框架BeamTransFuser,利用跨模态相关性提升波束预测性能,开发剪枝方案减少延迟,引入生成模型处理模态缺失,实验证明该方案优于现有基线。

Comments 15 pages, 5 figures

Journal ref IEEE Transactions on Mobile Computing, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31093 2026-07-01 cs.DC 新提交 82%

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference

Omni-Flow:面向多模态推理的统一工作流编排与分布式KV缓存共享框架

Bin Xiao, Jingfu Dong, Changran Wang, Yitian Chen, Xiaoyu Zhao, Yuqi Peng, Jianping Lin, Yuchen Xie

专题命中 多模态生成 :multimodal(title,abstract);omni-modal(abstract)

AI总结 提出Omni-Flow框架,通过三层抽象(控制流、数据流、计算流)统一编排多模态推理工作流,实现异构单元协同、分布式KV缓存共享与高效传输,支持多种多模态场景。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19534 2026-06-19 cs.CV cs.AI cs.CL 新提交 82%

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

PerceptionDLM:基于多模态扩散语言模型的并行区域感知

Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, Tao Zhang, Jacky Mai, Yihan Wang, Haochen Wang, Jinbin Bai, Ling Yang, Yunhai Tong

机构 * Peking University(北京大学) MSALab ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出PerceptionDLM,利用扩散语言模型的并行解码特性,通过高效提示和结构化注意力掩码实现多区域并行感知,显著提升推理效率,并构建ParaDLC-Bench基准进行评估。

Comments Code available at https://github.com/MSALab-PKU/PerceptionDLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10340 2026-06-10 cs.RO 新提交 82%

OMG: Omni-Modal Motion Generation for Generalist Humanoid Control

OMG: 面向通用人形机器人的全模态运动生成

Siqiao Huang, Kun-Ying Lee, Dongming Qiao, Guanqi He, Zhenyu Wang, Yitang Li, Shaoting Zhu, Hang Zhao

机构 * Tsinghua University(清华大学)

专题命中 多模态生成 :omni-modal(title,abstract);multi-modal(abstract)

AI总结 提出OMG框架,通过精心策划的数据流程和扩散模型,实现基于语言、音频和参考动作的全模态全身控制,展示了最先进的性能和可扩展性。

Comments Project Page: https://tsinghua-mars-lab.github.io/OMG/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09574 2026-05-29 cs.CV cs.AI cs.CL 82%

MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

MENTOR: 面向自回归视觉生成模型的高效多模态条件微调

Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen, Jiuxiang Gu, Wen Xiao, Minjia Zhang, Junjie Hu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Tsinghua University(清华大学) Peking University(北京大学) Microsoft(微软公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出MENTOR框架,通过两阶段训练范式实现自回归图像生成器与多模态输入的细粒度token级对齐,无需辅助适配器或交叉注意力模块,在DreamBench++上取得优异性能。

Comments Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16527 2026-05-19 cs.LG cs.AI cs.CL cs.CV 82%

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

超越表面遗忘:多模态大语言模型中Hallucinations的锐度感知鲁棒擦除

Xianya Fang, Feiyang Ren, Xiang Chen, Yu Tian, Zhen Bi, Haiyang Yu, Sheng-Jun Huang

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) Institute for AI, Tsinghua University(清华大学人工智能研究院) Huzhou University(湖州大学) Institute of Dataspace, Hefei Comprehensive National Science Center(合肥综合性国家科学中心数据空间研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出SARE方法,通过目标导向的min-max优化和Targeted-SAM机制,解决多模态大语言模型中 hallucinations 的鲁棒擦除问题,提升模型稳定性与擦除效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02494 2026-04-30 cs.CL cs.AI cs.CV 82%

MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text

MINOS:一种用于图像与文本双向生成的多模态评估模型

Junzhe Zhang, Huixuan Zhang, Xinyu Hu, Li Lin, Mingqi Gao, Shi Qiu, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王萱计算机技术研究所) School of Physics, Peking University(北京大学物理学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MINOS模型,通过严格质量控制策略构建Minos-57K多模态评估数据集,结合SFT和偏好对齐训练策略,在16个跨领域数据集中取得最佳性能。

Comments Accepted to the Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20434 2026-04-23 cs.IR 82%

Discrete Preference Learning for Personalized Multimodal Generation

离散偏好学习用于个性化多模态生成

Yuting Zhang, Ying Sun, Dazhong Shen, Ziwei Xie, Feng Liu, Changwang Zhang, Xiang Liu, Jun Wang, Hui Xiong

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出DPPMG框架,通过离散偏好模型学习多模态偏好并注入生成器,解决连续偏好与离散输入的矛盾及模态不一致问题,实验证明生成的多模态内容一致且个性化。

Comments be accepted to SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18871 2026-04-22 cs.CV cs.AI cs.CL 82%

OmniGen2: Towards Instruction-Aligned Multimodal Generation

OmniGen2:面向指令对齐的多模态生成

Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao, Xin Luo, Yueze Wang, Wanli Li, Xiyan Jiang, Yexin Liu, Junjie Zhou, Ze Liu, Ziyi Xia, Chaofan Li, Haoge Deng, Jiahao Wang, Kun Luo, Bo Zhang, Defu Lian, Xinlong Wang, Zhongyuan Wang, Tiejun Huang, Zheng Liu

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 OmniGen2通过双解码路径和解耦图像分词器,实现文本和图像生成的统一解决方案,提升多模态生成任务的性能与一致性,同时提供开源模型和数据集支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15420 2026-04-14 cs.LG 82%

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

FlowBind: 任意到任意生成的高效方法

Yeonwoo Cha, Semin Kim, Jinhyeon Kwon, Seunghoon Hong

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :any-to-any(title,abstract);cross-modal(abstract)

AI总结 本文提出FlowBind框架,通过共享潜在空间和双向流实现高效任意模态生成,减少数据和计算成本,实验显示生成质量与传统方法相当。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13130 2026-04-07 cs.CV cs.AI cs.CL 82%

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

ZINA:多模态细粒度幻觉检测与编辑

Yuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham Neubig

机构 * Keio AI Research Center(庆应义塾大学人工智能研究中心) Keio University(庆应义塾大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出ZINA方法,用于多模态大语言模型的细粒度幻觉检测与编辑,通过构建VisionHall数据集验证了其在检测和编辑任务中的优越性。

Comments CVPR 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27135 2026-03-31 cs.LG stat.ML 82%

Spectral-Aware Text-to-Time Series Generation with Billion-Scale Multimodal Meteorological Data

具有十亿级多模态气象数据的频谱感知文本到时间序列生成

Shijie Zhang

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院,沈阳,中国)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出MTransformer模型,通过频谱提示生成器实现文本到时间序列的精确语义控制,利用大规模气象数据提升生成质量与跨模态对齐精度。

Comments Accepted By IJCNN 2026 (WCCI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10343 2026-03-12 eess.SP 82%

Multi-Modal Intelligent Channel Modeling: From Fine-tuned LLMs to Pre-trained Foundation Models

多模态智能信道建模:从微调大语言模型到预训练基础模型

Lu Bai, Zengrui Han, Mingran Sun, Xiang Cheng

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 本文提出多模态智能信道建模,通过微调大语言模型和预训练基础模型,实现6G无线通信的精确预测与灵活建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06508 2026-03-09 cs.LG 82%

When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models

当一种模态统治一切:多模态扩散模型中的后门模态崩溃

Qitong Wang, Haoran Dai, Haotian Zhang, Christopher Rasmussen, Binghui Wang

机构 * University of Delaware(德克萨斯大学) Illinois Institute of Technology(伊利诺伊理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文研究了多模态扩散模型中后门攻击的模态崩溃现象,发现攻击常依赖单一模态,跨模态交互反而降低攻击效果,揭示了现有评估的盲点。

Comments Accepted to the ICLR 2026 Workshop on Principled Design for Trustworthy AI. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02273 2026-03-04 cs.LG 82%

Graph Attention Based Prioritization of Disease Responsible Genes from Multimodal Alzheimer's Network

基于图注意力的多模态阿尔茨海默病相关基因优先级排序

Binon Teji, Subhajit Bandyopadhyay, Swarup Roy

机构 * Network Reconstruction & Analysis Lab and Department of Computer Applications, Sikkim University(网络重建与分析实验室和计算机应用系,西西米大学) Network Reconstruction & Analysis Lab, Department of Computer Applications, Sikkim University(网络重建与分析实验室,计算机应用系,西西米大学) Department of Computer Science & Engineering, Tezpur University(计算机科学与工程系,泰朱大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 NETRA通过多模态图变换器框架,利用注意力机制优先排序阿尔茨海默病相关基因,显著提升疾病通路富集分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01965 2026-03-03 cs.LG q-bio.QM 82%

CoVAE: correlated multimodal generative modeling

CoVAE:相关多模态生成建模

Federico Caretti, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati(国际先进研究学院)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 CoVAE通过捕捉多模态之间的相关性,改进了多模态生成建模,实现了准确的跨模态重建和不确定性量化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15281 2026-03-03 cs.IR cs.LG 82%

MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation

MMQ: 多模态混合量化标记化用于语义ID生成和用户行为适应

Yi Xu, Moyu Zhang, Chenxuan Li, Zhihao Liao, Haibo Xing, Hao Deng, Jinxin Hu, Yu Zhang, Xiaoyi Zeng, Jing Zhang

机构 * Alibaba Group(阿里巴巴集团) Peking University(北京大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Wuhan University, School of Computer Science(武汉大学计算机学院)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 MMQ通过多模态混合量化标记化方法,解决推荐系统中多模态协同、特定性和行为适应的挑战,提升语义ID生成和用户行为适应能力。

Journal ref Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23366 2026-03-02 cs.HC cs.IR 82%

Doc To The Future: Infomorphs for Interactive, Multimodal Document Transformation and Generation

文档面向未来:用于交互式、多模态文档转换与生成的Infomorphs

Balasaravanan Thoravi Kumaravel

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Doc To The Future提出Infomorphs,通过模块化、用户可控的AI增强转换,实现交互式多模态文档转换与生成,提升生成式AI在信息工作中的透明度和模块化交互。

详情

展开后加载摘要…

URL PDF HTML 收藏