arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

1911.06663 2019-11-18 cs.LG cs.CV stat.ML 79%

MMGAN: Generative Adversarial Networks for Multi-Modal Distributions

Teodora Pandeva, Matthias Schubert

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04888 2019-06-13 cs.RO cs.CV cs.SY eess.SY 79%

Adaptive Navigation Scheme for Optimal Deep-Sea Localization Using Multimodal Perception Cues

Arturo Gomez Chavez, Qingwen Xu, Christian A. Mueller, Sören Schwertfeger, Andreas Birk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to IROS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13443 2019-06-03 cs.CL 79%

Symbol Emergence as an Interpersonal Multimodal Categorization

Yoshinobu Hagiwara, Hiroyoshi Kobayashi, Akira Taniguchi, Tadahiro Taniguchi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.13304 2019-04-25 cs.CV 79%

Acute and sub-acute stroke lesion segmentation from multimodal MRI

Albert Clèrigues, Sergi Valverde, Jose Bernal, Jordi Freixenet, Arnau Oliver, Xavier Lladó

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01735 2019-04-04 cs.CL 79%

Multi-Modal Generative Adversarial Network for Short Product Title Generation in Mobile E-Commerce

Jian-Guo Zhang, Pengcheng Zou, Zhao Li, Yao Wan, Xiuming Pan, Yu Gong, Philip S. Yu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by NAACL-HLT 2019. arXiv admin note: substantial text overlap with arXiv:1811.04498

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11955 2018-11-22 cs.CL 79%

Improving Context Modelling in Multimodal Dialogue Generation

Shubham Agarwal, Ondrej Dusek, Ioannis Konstas, Verena Rieser

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Journal ref Proceedings of the 11th International Conference on Natural Language Generation, pages 129-134, Tilburg, The Netherlands, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04498 2018-11-13 cs.CL 79%

Product Title Refinement via Multi-Modal Generative Adversarial Learning

Jianguo Zhang, Pengcheng Zou, Zhao Li, Yao Wan, Ye Liu, Xiuming Pan, Yu Gong, Philip S. Yu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Workshop on Visually Grounded Interaction and Language, NIPS, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03257 2018-05-10 cs.CL 79%

Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog

Jiaping Zhang, Tiancheng Zhao, Zhou Yu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00410 2018-04-03 cs.CV 79%

SyncGAN: Synchronize the Latent Space of Cross-modal Generative Adversarial Networks

Wen-Cheng Chen, Chien-Wen Chen, Min-Chun Hu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments 9 pages, Part of this work is accepted by IEEE International Conference on Multimedia Expo 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.11550 2018-04-02 cs.CV 79%

Multi-modal Disease Classification in Incomplete Datasets Using Geometric Matrix Completion

Gerome Vivar, Andreas Zwergal, Nassir Navab, Seyed-Ahmad Ahmadi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.06151 2017-09-20 cs.CV q-bio.NC q-bio.QM 79%

Multi-modal analysis of genetically-related subjects using SIFT descriptors in brain MRI

Kuldeep Kumar, Laurent Chauvin, Mathew Toews, Olivier Colliot, Christian Desrosiers

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Journal ref Proc. Computational Diffusion MRI, MICCAI Workshop, Québec City, Canada, September 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02337 2017-06-09 cs.CV cs.LG 79%

Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Network

Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, C. Lee Giles

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2017 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.02841 2017-04-11 cs.HC cs.CL 79%

From Modal to Multimodal Ambiguities: a Classification Approach

Maria Chiara Caschera, Fernando Ferri, Patrizia Grifoni

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 23 pages

Journal ref JNIT (Journal of Next Generation Information Technology), Volume 4 Issue 5, July, 2013,Pages 87-109, ISSN 2092-8637. GlobalCIS (Convergence Information Society, Republic of Korea)

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.02621 2017-04-11 cs.AI stat.ML 79%

Mixed Graphical Models for Causal Analysis of Multi-modal Variables

Andrew J Sedgewick, Joseph D. Ramsey, Peter Spirtes, Clark Glymour, Panayiotis V. Benos

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08472 2016-11-28 cs.CV 79%

Multimodal Latent Variable Analysis

Vardan Papyan, Ronen Talmon

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.01801 2016-08-29 cs.CV cs.LG stat.ML 79%

Variational methods for Conditional Multimodal Deep Learning

Gaurav Pandey, Ambedkar Dukkipati

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.07481 2016-06-27 cs.CL 79%

CUNI System for WMT16 Automatic Post-Editing and Multimodal Translation Tasks

Jindřich Libovický, Jindřich Helcl, Marek Tlustý, Pavel Pecina, Ondřej Bojar

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to the First Conference of Machine Translation (WMT16)

详情

展开后加载摘要…

URL PDF HTML 收藏
1003.1458 2010-03-09 cs.CR cs.CV 79%

Secured Cryptographic Key Generation From Multimodal Biometrics: Feature Level Fusion of Fingerprint and Iris

A. Jagadeesan, K. Duraiswamy

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Pages IEEE format, International Journal of Computer Science and Information Security, IJCSIS February 2010, ISSN 1947 5500, http://sites.google.com/site/ijcsis/

Journal ref International Journal of Computer Science and Information Security, IJCSIS, Vol. 7, No. 2, pp. 028-037, February 2010, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11136 2025-12-10 hep-ph hep-ex hep-th nucl-ex nucl-th 79%

Heavy-flavor multimodal fragmentation to $S$-wave pentacharms at next-generation hadron colliders

重味多模碎片化到下一代强子对撞机的S波五重态

Francesco Giovanni Celiberto

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 研究了在下一代强子对撞机中通过多模碎片化产生S波五重态的机制及现象学影响。

Comments 49 pages, 9 figures, 245 references, published in Eur. Phys. J. C. One novel set of multimodal (direct multicharm and diquark-like initial-scale inputs) "PentaQuarks with 5 heavy Quarks" (PQ5Q1.0) NLO collinear fragmentation functions released in LHAPDF format and publicly available from https://github.com/FGCeliberto/Collinear_FFs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21835 2025-10-28 cs.LG cs.AI cs.CL cs.CV 79%

A Multimodal, Multitask System for Generating E Commerce Text Listings from Images

Nayan Kumar Singh

专题命中 多模态生成 :multimodal(title,comments);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 10 figures, 11 tables. Code can be found at: https://github.com/SinghNayanKumar/multimodal-product-lister/

详情

展开后加载摘要…

URL PDF HTML 收藏
0708.3575 2009-12-01 cs.HC 79%

How really effective are Multimodal Hints in enhancing Visual Target Spotting? Some evidence from a usability study

Suzanne Kieffer, Noëlle Carbonell

专题命中 多模态生成 :multimodal(title,abstract)

Comments 9 pages

Journal ref Journal on Multimodal Interaction (JMUI), 1 (2007) 1-9

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18504 2026-07-27 cs.LG cs.AI cs.CV 版本更新 79%

Now We Know? A Systematic Comparison of TerraMind and THOR

我们现在知道了吗?TerraMind和THOR的系统比较

Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-Børre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci

机构 * University of Münster(明斯特大学) IBM Research(IBM研究院) Norwegian Computing Center(挪威计算中心)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.CV、cs.AI

AI总结 通过对TerraMind和THOR两个地理空间基础模型对比,研究其在补丁大小、解码器复杂性等方面的差异轴,发现架构设计选择对性能差异影响更大,体现互补投资策略,还得出假设和诊断消融方法,有望推广到未来模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08492 2026-07-02 cs.CV cs.AI 新提交 79%

Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation

眼见为实:基于视觉锚点的提示重写对齐用于文本到图像生成

Xuanyi Liu, Deyi Ji, Junyu Lu, Jing Wang, Lanyun Zhu, Qianxiong Xu, Xuhang Chen, Tianrun Chen, Siwei Ma

机构 * Peking University(北京大学) Tencent(腾讯) Dalian University of Technology(大连理工大学) Nanyang Technological University(南洋理工大学) University of Cambridge(剑桥大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出FaithRewriter框架,利用多模态大模型生成中间视觉线索,结合大语言模型生成视觉锚定的增强提示,再蒸馏至小模型,以缩小用户意图与生成图像之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30131 2026-05-29 cs.CL cs.CV 79%

CCS: Clinical Consensus Selection for Radiology Report Generation

CCS:放射学报告生成的临床共识选择

Xi Zhang, Yingshu Li, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算机科学学院) School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院) Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 提出CCS框架,通过采样多个候选报告并选择临床共识最高的一个,以改进放射学报告生成在推理时的质量。

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12480 2026-05-13 cs.CV cs.AI 79%

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

OmniNFT: 多模态联合音频视频生成的模态感知强化学习

Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) JD Explore Academy(京东探索研究院)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出OmniNFT,通过模态感知强化学习框架解决多模态联合生成中的模态一致性、梯度不平衡和信用分配问题,提升音频视频感知质量与同步性能。

Comments Project page: https://zghhui.github.io/OmniNFT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12138 2026-05-13 cs.CV cs.CL cs.IR 79%

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

为广告设计:基于统一自回归模型的个性化广告图像和文本生成

Yexing Xu, Wei Feng, Shen Zhang, Haohan Wang, Yuxin Qin, Yaoyu Li, Ao Ma, Yuhao Luo, Lu Wang, Xudong Ren, Haoran Wang, Run Ling, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Longguang Wang, Yulan Guo

机构 * Sun Yat-Sen University(中山大学) Northeastern University(东北大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

AI总结 本文提出统一广告生成模型Uni-AdGen,通过单个自回归框架生成个性化图文广告,结合前景感知模块和指令微调提升生成质量,并引入大规模广告数据集和新指标提升个性化生成效果。

Comments 22 pages, 19 figures, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18269 2026-04-21 cs.CL cs.CV 79%

TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation

TextTIGER:基于实体提示精炼的文本智能生成用于文本到图像生成

Shintaro Ozaki, Tomoyuki Jinno, Kazuki Hayashi, Yusuke Sakai, Jingun Kwon, Hidetaka Kamigaito, Katsuhiko Hayashi, Manabu Okumura, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技術研究所) The University of Tokyo(东京大学) Chungnam National University(Chungnam国立大学) Institute of Science Tokyo(东京科学研究所)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 TextTIGER通过增强外部信息和大语言模型总结,优化实体描述以提升文本到图像生成性能,实验表明其在多种评估指标上优于仅使用描述的提示方法。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00007 2026-04-02 cs.CL cs.AI 79%

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

Dynin-Omni: 多模态统一的大扩散语言模型

Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee, Sieun Hyeon, Mintaek Lim, Yunseok Han, Dogeun Kim, Hoeun Lee, Hyunggeun Kim, Jaeyoung Do

机构 * AIDAS Lab Seoul National University [1.5ex] Project Page Code Model Demo -2em

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.CL、cs.AI

AI总结 Dynin-Omni首次提出基于掩码扩散的多模态基础模型,统一文本、图像、语音的理解与生成,以及视频理解,通过共享离散令牌空间实现双向上下文迭代优化。

Comments Project Page: https://dynin.ai/omni/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01019 2025-04-22 cs.CV cs.AI 79%

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

Ziyang Zhang, Yang Yu, Yucheng Chen, Xulei Yang, Si Yong Yeo

机构 * MedVisAI Lab(MedVisAI实验室) ECE, Northwestern University(电子工程系,西北大学) Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究所(I2R),A*STAR,新加坡) Lee Kong Chian School of Medicine, Nanyang Technological University(Lee Kong Chian医学院,南洋理工大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments To be pubilshed in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00289 2025-04-03 cs.CV cs.AI cs.LG 79%

Dual Diffusion for Unified Image Generation and Understanding

Zijie Li, Henry Li, Yichun Shi, Amir Barati Farimani, Yuval Kluger, Linjie Yang, Peng Wang

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏