arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2309.15739 2023-09-28 cs.CL cs.AI 81%

Experience and Evidence are the eyes of an excellent summarizer! Towards Knowledge Infused Multi-modal Clinical Conversation Summarization

Abhisek Tiwari, Anisha Saha, Sriparna Saha, Pushpak Bhattacharyya, Minakshi Dhar

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15142 2023-08-30 cs.CV cs.AI q-bio.NC 81%

A Multimodal Visual Encoding Model Aided by Introducing Verbal Semantic Information

Shuxiao Ma, Linyuan Wang, Bin Yan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10910 2023-08-23 eess.IV cs.AI cs.CV 81%

Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI Reconstruction

Yunlu Yan, Chun-Mei Feng, Yuexiang Li, Rick Siow Mong Goh, Lei Zhu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00400 2023-08-03 cs.CL cs.MM 81%

ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue Generation

Bo Zhang, Jian Wang, Hui Ma, Bo Xu, Hongfei Lin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments ACM Multimedia 2023 Accpeted, Repo: https://github.com/zhangbo-nlp/ZRIGF

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08597 2023-07-18 cs.CV cs.CL cs.RO 81%

Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions

Yui Iioka, Yu Yoshida, Yuiga Wada, Shumpei Hatanaka, Komei Sugiura

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted for presentation at IROS2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04643 2023-07-11 cs.CL cs.AI 81%

MultiQG-TI: Towards Question Generation from Multi-modal Sources

Zichao Wang, Richard Baraniuk

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at BEA workshop 2023; code https://github.com/moonlightlane/MultiQG-TI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16650 2023-06-30 cs.CL cs.AI 81%

Multi-source Semantic Graph-based Multimodal Sarcasm Explanation Generation

Liqiang Jing, Xuemeng Song, Kun Ouyang, Mengzhao Jia, Liqiang Nie

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2023 main conference

Journal ref ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19012 2023-06-01 cs.CV cs.AI 81%

StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation

Chi Zhang, Yiwen Chen, Yijun Fu, Zhenglin Zhou, Gang YU, Billzb Wang, Bin Fu, Tao Chen, Guosheng Lin, Chunhua Shen

专题命中 多模态生成 :image-text(title,abstract);分类 cs.CV、cs.AI

Comments Project page: https://github.com/icoz69/StyleAvatar3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17433 2023-05-30 cs.CV cs.CL 81%

A Unified Framework for Slot based Response Generation in a Multimodal Dialogue System

Mauajama Firdaus, Avinash Madasu, Asif Ekbal

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published in the journal Multimedia Tools and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09990 2023-05-18 cs.CL cs.MM 81%

Dual Semantic Knowledge Composed Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie, Tat-Seng Chua

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments SIGIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13855 2023-04-28 cs.CV cs.AI cs.CY cs.LG 81%

Multimodal Composite Association Score: Measuring Gender Bias in Generative Multimodal Models

Abhishek Mandal, Susan Leavy, Suzanne Little

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution has been accepted at the Fourth International Workshop on Algorithmic Bias in Search and Recommendation held as a part of the 45th European Conference on Information Retrieval (ECIR 2023) and will be published soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.14579 2023-04-12 cs.CL cs.CV cs.LG 81%

Competence-based Multimodal Curriculum Learning for Medical Report Generation

Fenglin Liu, Shen Ge, Yuexian Zou, Xian Wu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by ACL 2021 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12824 2023-03-23 cs.CV cs.CL 81%

Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation

Tsu-Jui Fu, Licheng Yu, Ning Zhang, Cheng-Yang Fu, Jong-Chyi Su, William Yang Wang, Sean Bell

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR'23

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10305 2023-02-22 cs.CV cs.AI cs.LG 81%

Analyzing Multimodal Objectives Through the Lens of Generative Diffusion Guidance

Chaerin Kong, Nojun Kwak

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15461 2022-11-29 cs.CL cs.AI 81%

LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation

Hongcheng Guo, Jiaheng Liu, Haoyang Huang, Jian Yang, Zhoujun Li, Dongdong Zhang, Zheng Cui, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07210 2022-11-15 cs.CV cs.AI 81%

Grafting Pre-trained Models for Multimodal Headline Generation

Lingfeng Qiao, Chen Wu, Ye Liu, Haoyuan Peng, Di Yin, Bo Ren

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12121 2022-07-26 cs.SD cs.CV cs.GR eess.AS 81%

Cross-Modal Contrastive Representation Learning for Audio-to-Image Generation

HaeChun Chung, JooYong Shim, Jong-Kook Kim

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments 7 pages, 3 figures, Accepted to MUE 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04818 2022-07-12 cs.CV cs.CL 81%

Cross-modal Prototype Driven Network for Radiology Report Generation

Jun Wang, Abhir Bhalerao, Yulan He

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01823 2022-07-06 cs.CL cs.CV 81%

Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation

Bin Li, Yixuan Weng, Ziyu Ma, Bin Sun, Shutao Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted in NLPCC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.15011 2022-06-02 eess.IV cs.CL cs.CV 81%

Radiology Report Generation with a Learned Knowledge Base and Multi-modal Alignment

Shuxin Yang, Xian Wu, Shen Ge, S. Kevin Zhou, Li Xiao

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04006 2022-05-10 cs.CL cs.AI 81%

Data Augmentation with Paraphrase Generation and Entity Extraction for Multimodal Dialogue System

Eda Okur, Saurav Sahay, Lama Nachman

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Proceedings of the 13th International Conference on Language Resources and Evaluation (LREC 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09081 2022-03-31 cs.CV cs.AI cs.RO 81%

CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic Data

Qi Yan, Jianhao Zheng, Simon Reding, Shanci Li, Iordan Doytchinov

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.AI

Comments CVPR 2022. Project page: https://crossloc.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05328 2021-12-21 cs.CL cs.AI 81%

Multimodal Interactions Using Pretrained Unimodal Models for SIMMC 2.0

Joosung Lee, Kijong Han

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to DSTC10 challenge wokrshop at AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05692 2021-12-13 cs.CV cs.AI cs.HC cs.LG 81%

VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling

Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani, Alexey Gritsenko

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01682 2021-08-09 cs.CL cs.CV 81%

Exploiting BERT For Multimodal Target Sentiment Classification Through Input Space Translation

Zaid Khan, Yun Fu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ACM Multimedia 2021 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04806 2021-07-15 cs.SD cs.CV eess.AS eess.IV 81%

Speech2Video: Cross-Modal Distillation for Speech to Video Generation

Shijing Si, Jianzong Wang, Xiaoyang Qu, Ning Cheng, Wenqi Wei, Xinghua Zhu, Jing Xiao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Accepted by InterSpeech2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05468 2021-07-13 cs.CV cs.AI cs.RO 81%

Visual-Tactile Cross-Modal Data Generation using Residue-Fusion GAN with Feature-Matching and Perceptual Losses

Shaoyu Cai, Kening Zhu, Yuki Ban, Takuji Narumi

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures, Accepted by IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13135 2021-05-28 cs.CL cs.AI cs.LG 81%

Self-Supervised Multimodal Opinion Summarization

Jinbae Im, Moonki Kim, Hoyeop Lee, Hyunsouk Cho, Sehee Chung

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10299 2021-04-22 cs.GR cs.CV cs.LG cs.SD eess.AS 81%

Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices

Cho-Ying Wu, Ke Xu, Chin-Cheng Hsu, Ulrich Neumann

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、eess.AS

Comments Project page: https://choyingw.github.io/works/Voice2Mesh/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.10832 2020-12-17 cs.CL cs.CV cs.LG 81%

What BERT Sees: Cross-Modal Transfer for Visual Question Generation

Thomas Scialom, Patrick Bordes, Paul-Alexis Dray, Jacopo Staiano, Patrick Gallinari

专题命中 多模态生成 :cross-modal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments INLG 2020

详情

展开后加载摘要…

URL PDF HTML 收藏