arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4932 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4932 篇

2311.11090 2023-11-21 cs.CV cs.AI cs.CL cs.LG 82%

Beyond Images: An Integrative Multi-modal Approach to Chest X-Ray Report Generation

Nurbanu Aksoy, Serge Sharoff, Selcuk Baser, Nishant Ravikumar, Alejandro F Frangi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17842 2023-10-31 cs.CV cs.CL cs.MM 82%

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin Murphy, Alexander G. Hauptmann, Lu Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments NeurIPS 2023 spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10765 2023-10-24 cs.CV cs.AI cs.CL 82%

BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys

Yu Gu, Jianwei Yang, Naoto Usuyama, Chunyuan Li, Sheng Zhang, Matthew P. Lungren, Jianfeng Gao, Hoifung Poon

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page & demo: https://aka.ms/biomedjourney

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13361 2023-10-23 cs.CV cs.AI cs.CL 82%

Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation

Wenyu Guo, Qingkai Fang, Dong Yu, Yang Feng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2023 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02051 2023-08-24 cs.CV cs.AI cs.MM 82%

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10843 2023-08-22 cs.MM cs.CV cs.LG cs.SD eess.AS 82%

TranSTYLer: Multimodal Behavioral Style Transfer for Facial and Body Gestures Generation

Mireille Fares, Catherine Pelachaud, Nicolas Obin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11846 2023-05-22 cs.CV cs.CL cs.LG cs.SD eess.AS 82%

Any-to-Any Generation via Composable Diffusion

Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, Mohit Bansal

专题命中 多模态生成 :any-to-any(title);multimodal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments Project Page: https://codi-gen.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11832 2023-05-22 stat.ML cs.LG 82%

Improving Multimodal Joint Variational Autoencoders through Normalizing Flows and Correlation Analysis

Agathe Senellart, Clément Chadebec, Stéphanie Allassonnière

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06756 2023-03-31 cs.CV cs.AI cs.MM cs.NE 82%

Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features

Changde Du, Kaicheng Fu, Jinpeng Li, Huiguang He

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08672 2023-02-20 cs.LG cs.AI cs.CL cs.CV 82%

Multimodal Subtask Graph Generation from Instructional Videos

Yunseok Jang, Sungryull Sohn, Lajanugen Logeswaran, Tiange Luo, Moontae Lee, Honglak Lee

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13235 2022-11-28 cs.CL cs.AI cs.CV 82%

Unified Multimodal Model with Unlikelihood Training for Visual Dialog

Zihao Wang, Junli Wang, Changjun Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by the 30th ACM International Conference on Multimedia (ACM MM 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02127 2022-07-06 cs.LG stat.ML 82%

A survey of multimodal deep generative models

Masahiro Suzuki, Yutaka Matsuo

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments Published in Advanced Robotics

Journal ref Advanced Robotics, 36:5-6, 261-278, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11705 2022-05-25 cs.CV cs.AI cs.MM 82%

M6-Fashion: High-Fidelity Multi-modal Image Generation and Editing

Zhikang Li, Huiling Zhou, Shuai Bai, Peike Li, Chang Zhou, Hongxia Yang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments arXiv admin note: text overlap with arXiv:2105.14211

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03534 2022-05-10 cs.CL cs.CV cs.MM 82%

Attract me to Buy: Advertisement Copywriting Generation with Multimodal Multi-structured Information

Zhipeng Zhang, Xinglin Hou, Kai Niu, Zhongzhen Huang, Tiezheng Ge, Yuning Jiang, Qi Wu, Peng Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00590 2022-03-29 cs.CL cs.AI cs.CV cs.LG 82%

WebQA: Multihop and Multimodal QA

Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, Yonatan Bisk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR Camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09753 2021-10-20 cs.CV cs.CL cs.MM 82%

Unifying Multimodal Transformer for Bi-directional Image and Text Generation

Yupan Huang, Hongwei Xue, Bei Liu, Yutong Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACM MM 2021 (Industrial Track). Code: https://github.com/researchmm/generate-it

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08478 2021-09-20 cs.CL cs.CV cs.MM 82%

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li, Jie Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACL Fingdings 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04324 2021-09-17 cs.CL cs.AI cs.CV 82%

FairyTailor: A Multimodal Generative Framework for Storytelling

Eden Bensaid, Mauro Martino, Benjamin Hoover, Hendrik Strobelt

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments visit https://fairytailor.org/ and https://github.com/EdenBD/MultiModalStory-demo for web demo and source code

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09252 2020-05-20 cs.IR 82%

Multi-Modal Summary Generation using Multi-Objective Optimization

Anubhav Jangra, Sriparna Saha, Adam Jatowt, Mohammad Hasanuzzaman

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.01707 2020-01-07 cs.LG eess.IV stat.ML 82%

Meta-modal Information Flow: A Method for Capturing Multimodal Modular Disconnectivity in Schizophrenia

Haleh Falakshahi, Victor M. Vergara, Jingyu Liu, Daniel H. Mathalon, Judith M. Ford, James Voyvodic, Bryon A. Mueller, Aysenil Belger, Sarah McEwen, Steven G. Potkin, Adrian Preda, Hooman Rokham, Jing Sui, Jessica A. Turner, Sergey Plis, Vince D. Calhoun

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Journal ref IEEE Transactions on Biomedical Engineering, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03393 2019-11-11 stat.ML cs.LG 82%

Variational Mixture-of-Experts Autoencoders for Multi-Modal Deep Generative Models

Yuge Shi, N. Siddharth, Brooks Paige, Philip H. S. Torr

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03986 2019-10-18 cs.CL cs.AI cs.CV 82%

Multimodal Differential Network for Visual Question Generation

Badri N. Patro, Sandeep Kumar, Vinod K. Kurmi, Vinay P. Namboodiri

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2018 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.08129 2018-02-23 cs.AI cs.CL cs.CV 82%

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, Marcus Rohrbach

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:1612.04757

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17585 2026-04-21 cs.CV cs.AI cs.LG 82%

DGSSM: Diffusion guided state-space models for multimodal salient object detection

DGSSM:基于扩散引导的状态空间模型的多模态显著目标检测

Suklav Ghosh, Arijit Sur, Pinaki Mitra

机构 * Dept. of Computer Science and Engineering, Indian Institute of Technology, Guwahati(计算机科学与工程系,印度理工学院,果阿提)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出DGSSM,一种结合扩散模型结构先验和多尺度状态空间编码的多模态显著目标检测框架,通过迭代Mamba扩散细化机制提升边界精度,实验表明其在多个评估指标上优于现有方法。

Comments Accepted at ICPR 2026. Diffusion-guided Mamba framework for multimodal salient object detection. Evaluated on 13 benchmarks (RGB, RGB-D, RGB-T)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15512 2024-10-23 cs.CV cs.AI 82%

PixelBytes: Catching Unified Embedding for Multimodal Generation

Fabien Furfaro

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This article is an earlier version of my work arXiv:2410.01820 "PixelBytes: Catching Unified Representation for Multimodal Generation."

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04589 2024-01-05 cs.CL cs.AI 82%

TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models

Zhen Yang, Yingxue Zhang, Fandong Meng, Jie Zhou

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Multi-modal, Large Language Models, Tokenizer, Understanding and Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21138 2026-08-07 cs.CV 版本更新 81%

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

SEED: 用于可解释文本伪造检测的简单ViT与演化框架

Kahim Wong, Kemou Li, Yiming Chen, Haiwei Wu, Jiantao Zhou

机构 * State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系智慧城市物联网国家重点实验室) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

专题命中 多模态生成 :MLLM(summary_cn,abstract);分类 cs.CV

AI总结 提出SEED系统,结合相似性引导数据增强、单一ViT联合检测与定位、以及基于MLLM的演化报告生成,在ACM MM 2026文本伪造挑战赛中获得第三名。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24157 2026-07-28 cs.CV 新提交 81%

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

UniGen-AR:通过自回归建模统一视觉生成

Zhipeng Bao, Zhen Zhu, Nupur Kumari, Anurag Bagchi, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳 - 香槟分校) Toyota Research Institute(丰田研究院)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 研究统一视觉生成问题,提出UniGen-AR框架,将通用多模态语言模型与视觉自回归解码器配对,能为多任务生成图像值输出,相比基于扩散的基线,推理延迟低达19倍,确立视觉自回归建模为统一视觉生成的高效主干。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01050 2026-07-07 cs.CV 新提交 81%

GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

GeoSearcher: 基于锚点引导的渐进推理遥感视觉定位与过程监督

Dianyu Wang, Peirong Zhang, Xuyang Li, Xiaoxuan Liu, Lei Wang

机构 * Key Laboratory of Target Cognition and Application Technology (TCAT), Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出GeoSearcher,通过锚点引导的渐进推理和过程监督,将遥感视觉定位转化为两阶段过程,解决小目标定位和复杂查询的挑战。

Comments 14 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20273 2026-05-21 cs.LG cs.AI 81%

Modality-Decoupled Online Recursive Editing

模态解耦的在线递归编辑

Siyuan Li, Youyuan Zhang, Fangming Liu, Jing Li

机构 * Harbin Institute of Technology, Shenzhen, China.(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, China.(鹏城实验室) Huazhong University of Science and Technology, China(华中科技大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出M-ORE,一种用于持续多模态大语言模型适应的模态解耦在线递归编辑器,通过统一的近端投影公式和Sherman-Morrison递归实现常数级的每编辑开销,从而在保持模块局部统计信息和固定正交低秩编辑子空间的同时,减少长周期干扰,提升可靠性、通用性和局部性。

详情

展开后加载摘要…

URL PDF HTML 收藏