arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2207.04713 2022-07-12 cs.CL 79%

GMN: Generative Multi-modal Network for Practical Document Information Extraction

Haoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to NAACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08970 2022-06-22 cs.CV 79%

MultiEarth 2022 -- The Champion Solution for the Matrix Completion Challenge via Multimodal Regression and Generation

Bo Peng, Hongchen Liu, Hang Zhou, Yuchuan Gou, Jui-Hsin Lai

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2022, MultiEarth 2022, Matrix Completion Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09808 2022-06-15 cs.CV 79%

Integrated Construction of Multimodal Atlases with Structural Connectomes in the Space of Riemannian Metrics

Kristen M. Campbell, Haocheng Dai, Zhe Su, Martin Bauer, P. Thomas Fletcher, Sarang C. Joshi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://www.melba-journal.org/papers/2022:016.html. arXiv admin note: substantial text overlap with arXiv:2103.05730

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13871 2022-06-14 cs.SD cs.GR cs.LG eess.AS 79%

Transflower: probabilistic autoregressive dance generation with multimodal attention

Guillermo Valle-Pérez, Gustav Eje Henter, Jonas Beskow, André Holzapfel, Pierre-Yves Oudeyer, Simon Alexanderson

专题命中 多模态生成 :multimodal(title,abstract);分类 eess.AS

Comments Article presented at SIGGRAPH Asia 2021, and published in ACM Transactions on Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05039 2022-06-13 cs.CV 79%

Image Generation with Multimodal Priors using Denoising Diffusion Probabilistic Models

Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M Patel

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01988 2022-06-07 cs.CV 79%

Cross-modal Clinical Graph Transformer for Ophthalmic Report Generation

Mingjie Li, Wenjia Cai, Karin Verspoor, Shirui Pan, Xiaodan Liang, Xiaojun Chang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13258 2022-04-29 cs.CL 79%

Cross-modal Memory Networks for Radiology Report Generation

Zhihong Chen, Yaling Shen, Yan Song, Xiang Wan

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL

Comments Natural Language Processing. 11 pages, 6 figures. ACL-IJCNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12667 2022-04-28 cs.CV 79%

MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation

Inkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Sparsh Garg, In So Kweon, Kuk-Jin Yoon

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14211 2022-02-22 cs.CV 79%

M6-UFC: Unifying Multi-Modal Controls for Conditional Image Synthesis via Non-Autoregressive Generative Transformers

Zhu Zhang, Jianxin Ma, Chang Zhou, Rui Men, Zhikang Li, Ming Ding, Jie Tang, Jingren Zhou, Hongxia Yang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS21

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03608 2022-02-01 cs.LG cs.AI 79%

How to Sense the World: Leveraging Hierarchy in Multimodal Perception for Robust Reinforcement Learning Agents

Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at the International Conference on Autonomous Agents and MultiAgent Systems (AAMAS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09702 2021-10-22 cs.CL 79%

A non-hierarchical attention network with modality dropout for textual response generation in multimodal dialogue systems

Rongyi Sun, Borun Chen, Qingyu Zhou, Yinghui Li, YunBo Cao, Hai-Tao Zheng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to ICASSP2022 (currently under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02401 2021-10-12 cs.CL 79%

Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization

Tiezheng Yu, Wenliang Dai, Zihan Liu, Pascale Fung

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Long Paper Accepted in EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01229 2021-09-06 cs.CL cs.LG 79%

Multimodal Conditionality for Natural Language Generation

Michael Sollami, Aashish Jain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09953 2021-07-22 cs.CV 79%

Characterization Multimodal Connectivity of Brain Network by Hypergraph GAN for Alzheimer's Disease Analysis

Junren Pan, Baiying Lei, Yanyan Shen, Yong Liu, Zhiguang Feng, Shuqiang Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.08673 2021-07-20 eess.IV cs.CV cs.LG 79%

Input Agnostic Deep Learning for Alzheimer's Disease Classification Using Multimodal MRI Images

Aidana Massalimova, Huseyin Atakan Varol

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 4 pages, submitted to EMBC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00419 2021-07-19 cs.CL 79%

KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense Generation

Yiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer, Yunpu Ma, Roger Wattenhofer

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments ACL-IJCNLP 2021 main conference. The first three authors contribute equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.15309 2021-06-30 cs.CV 79%

Multimodal Semantic Scene Graphs for Holistic Modeling of Surgical Procedures

Ege Özsoy, Evin Pınar Örnek, Ulrich Eck, Federico Tombari, Nassir Navab

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14445 2021-06-01 cs.CL 79%

Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation

Shuhe Wang, Yuxian Meng, Xiaofei Sun, Fei Wu, Rongbin Ouyang, Rui Yan, Tianwei Zhang, Jiwei Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:2012.15015

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07854 2021-03-23 cs.CV 79%

Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and Synthesis

Jianhua Sun, Yuxuan Li, Hao-Shu Fang, Cewu Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10139 2021-03-19 cs.CV 79%

Learning Multimodal Affinities for Textual Editing in Images

Or Perel, Oron Anschel, Omri Ben-Eliezer, Shai Mazor, Hadar Averbuch-Elor

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ACM Transactions on Graphics 2021, to be presented in SIGGRAPH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12338 2021-02-01 cs.RO cs.AI 79%

Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Ting Han, Sina Zarrieß

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments The 2nd Workshop on NLG for HRI colocated with The 13th International Conference on Natural Language Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04776 2020-12-10 cs.LG cs.AI 79%

A Data-Driven Analytical Framework of Estimating Multimodal Travel Demand Patterns using Mobile Device Location Data

Chenfeng Xiong, Aref Darzi, Yixuan Pan, Sepehr Ghader, Lei Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09067 2020-10-20 cs.CV 79%

Multimodal semantic forecasting based on conditional generation of future features

Kristijan Fugošić, Josip Šarić, Siniša Šegvić

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to German Conference on Pattern Recognition 2020. 24 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12429 2020-04-28 cs.CL 79%

Towards Multimodal Response Generation with Exemplar Augmentation and Curriculum Optimization

Zeyang Lei, Zekang Li, Jinchao Zhang, Fandong Meng, Yang Feng, Yujiu Yang, Cheng Niu, Jie Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00431 2020-03-03 cs.AI 79%

A Study on Multimodal and Interactive Explanations for Visual Question Answering

Kamran Alipour, Jurgen P. Schulze, Yi Yao, Avi Ziskind, Giedrius Burachas

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments http://ceur-ws.org/Vol-2560/paper44.pdf

Journal ref Proceedings of the Workshop on Artificial Intelligence Safety (SafeAI 2020) co-located with 34th AAAI Conference on Artificial Intelligence (AAAI 2020), New York, USA, Feb 7, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06484 2020-02-18 cs.CL 79%

A Multimodal Dialogue System for Conversational Image Editing

Tzu-Hsiang Lin, Trung Bui, Doo Soon Kim, Jean Oh

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at 2nd Conversational AI Workshop at NeurIPS 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.09792 2019-11-20 cs.CL 79%

A Multi-Modal Chinese Poetry Generation Model

Dayiheng Liu, Quan Guo, Wubo Li, Jiancheng Lv

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06663 2019-11-18 cs.LG cs.CV stat.ML 79%

MMGAN: Generative Adversarial Networks for Multi-Modal Distributions

Teodora Pandeva, Matthias Schubert

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04888 2019-06-13 cs.RO cs.CV cs.SY eess.SY 79%

Adaptive Navigation Scheme for Optimal Deep-Sea Localization Using Multimodal Perception Cues

Arturo Gomez Chavez, Qingwen Xu, Christian A. Mueller, Sören Schwertfeger, Andreas Birk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to IROS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13443 2019-06-03 cs.CL 79%

Symbol Emergence as an Interpersonal Multimodal Categorization

Yoshinobu Hagiwara, Hiroyoshi Kobayashi, Akira Taniguchi, Tadahiro Taniguchi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏