arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

1905.12681 2020-04-06 cs.CV cs.LG 79%

What Makes Training Multi-Modal Classification Networks Hard?

Weiyao Wang, Du Tran, Matt Feiszli

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.09691 2020-03-20 cs.CV 79%

Multi-Modal Domain Adaptation for Fine-Grained Action Recognition

Jonathan Munro, Dima Damen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2020 for an oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09844 2020-03-13 cs.CV 79%

Adversarial Multimodal Network for Movie Question Answering

Zhaoquan Yuan, Siyuan Sun, Lixin Duan, Xiao Wu, Changsheng Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments We will revise the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06993 2020-03-10 cs.CV cs.RO 79%

Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations

Rogerio Bonatti, Ratnesh Madaan, Vibhav Vineet, Sebastian Scherer, Ashish Kapoor

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01043 2020-03-03 cs.CL cs.LG stat.ML 79%

Gated Mechanism for Attention Based Multimodal Sentiment Analysis

Ayush Kumar, Jithendra Vepa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to appear in ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12602 2020-02-27 cs.CV 79%

EV-Action: Electromyography-Vision Multi-Modal Action Dataset

Lichen Wang, Bin Sun, Joseph Robinson, Taotao Jing, Yun Fu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE International Conference on Automatic Face & Gesture Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01166 2020-02-26 cs.CL 79%

Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems

Hung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ACL 2019 (Long Paper)

Journal ref Association for Computational Linguistics (2019) 5612-5623

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.11657 2020-02-03 cs.CV 79%

Modality Compensation Network: Cross-Modal Adaptation for Action Recognition

Sijie Song, Jiaying Liu, Yanghao Li, Zongming Guo

专题命中 视频多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Trans. on Image Processing, 2020. Project page: http://39.96.165.147/Projects/MCN_tip2020_ssj/MCN_tip_2020_ssj.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.00628 2020-01-22 cs.CV cs.LG 79%

Sensor Fusion: Gated Recurrent Fusion to Learn Driving Behavior from Temporal Multimodal Data

Athma Narayanan, Avinash Siravuru, Behzad Dariush

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to Robotics and Automation Letters 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10982 2019-12-24 cs.CV 79%

DMCL: Distillation Multiple Choice Learning for Multimodal Action Recognition

Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.08291 2019-12-17 cs.CV 79%

Correlation Net: Spatiotemporal multimodal deep learning for action recognition

Novanto Yudistira, Takio Kurita

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Signal Processing: Image Communication, Volume 82, March 2020, 115731

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09826 2019-11-25 cs.LG cs.CL stat.ML 79%

Factorized Multimodal Transformer for Multimodal Sequential Learning

Amir Zadeh, Chengfeng Mao, Kelly Shi, Yiwei Zhang, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03974 2019-11-12 cs.MM cs.IR 79%

A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes

Pedro V. A. de Freitas, Paulo R. C. Mendes, Gabriel N. P. dos Santos, Antonio José G. Busson, Álan Livio Guedes, Sérgio Colcher, Ruy Luiz Milidiú

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00381 2019-11-04 cs.CV 79%

Multimodal Video-based Apparent Personality Recognition Using Long Short-Term Memory and Convolutional Neural Networks

Süleyman Aslan, Uğur Güdükbay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11482 2019-10-28 cs.LG cs.CV stat.ML 79%

Human Action Recognition Using Deep Multilevel Multimodal (M2) Fusion of Depth and Inertial Sensors

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04641 2019-10-11 cs.CV 79%

Cross-modal knowledge distillation for action recognition

Fida Mohammad Thoker, Juergen Gall

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Published in: 2019 IEEE International Conference on Image Processing (ICIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11604 2019-09-26 cs.AI 79%

An Extensible and Personalizable Multi-Modal Trip Planner

Xudong Liu, Christian Fritz, Matthew Klenk

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Published in the Proceedings of the 32nd International Florida Artificial Intelligence Research Society Conference, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01763 2019-09-05 cs.CV cs.LG 79%

Video Affective Effects Prediction with Multi-modal Fusion and Shot-Long Temporal Context

Jie Zhang, Yin Zhao, Longjun Cai, Chaoping Tu, Wu Wei

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08498 2019-08-23 cs.CV 79%

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima Damen

专题命中 视频多模态 :audio-visual(title);multi-modal(abstract);分类 cs.CV

Comments Accepted for presentation at ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.04955 2019-08-15 cs.RO cs.CV cs.HC cs.LG 79%

Probabilistic Multimodal Modeling for Human-Robot Interaction Tasks

Joseph Campbell, Simon Stepputtis, Heni Ben Amor

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project website: http://interactive-robotics.engineering.asu.edu/interaction-primitives Accompanying video: https://youtu.be/r5AqfxTDfLA

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13540 2019-06-03 cs.CV cs.LG stat.ML 79%

Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering

Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, Chang D. Yoo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IJCNN2019, oral

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.01263 2019-05-21 cs.CV 79%

Multimodal Explanations by Predicting Counterfactuality in Videos

Atsushi Kanehira, Kentaro Takemoto, Sho Inayoshi, Tatsuya Harada

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Camera ready version of CVPR'19

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01790 2019-05-14 cs.MM 79%

A multimodal lossless coding method for skeletons in videos

Mingzhou Liu, Xiaoyi He, Weiyao Lin, Xintong Han, Yanmin Zhu, Hongtao Lu, Hongkai Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments This manuscript is the accepted version for ICMEW (IEEE Intl. Conf. Multimedia & Expo Workshop), IEEE Intl. Conf. Multimedia & Expo Workshop (ICME), 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.03879 2019-04-30 cs.CV 79%

Cross and Learn: Cross-Modal Self-Supervision

Nawid Sayed, Biagio Brattoli, Björn Ommer

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments GCPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04357 2019-04-10 cs.CV 79%

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, Heng Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.04927 2018-11-28 cs.CV 79%

VIPL-HR: A Multi-modal Database for Pulse Estimation from Less-constrained Face Video

Xuesong Niu, Hu Han, Shiguang Shan, Xilin Chen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04595 2018-11-13 cs.CV 79%

Holistic Multi-modal Memory Network for Movie Question Answering

Anran Wang, Anh Tuan Luu, Chuan-Sheng Foo, Hongyuan Zhu, Yi Tay, Vijay Chandrasekhar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.10641 2018-07-30 cs.CV 79%

Multimodal Classification with Deep Convolutional-Recurrent Neural Networks for Electroencephalography

Chuanqi Tan, Fuchun Sun, Wenchang Zhang, Jianhua Chen, Chunfang Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

Journal ref Neural Information Processing. 2017:767-776

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10319 2018-06-28 cs.CV 79%

Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition

Dongliang He, Fu Li, Qijie Zhao, Xiang Long, Yi Fu, Shilei Wen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.02609 2018-06-08 cs.CV 79%

Learning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving

Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marin, David Martin, Lucio Marcenaro, Carlo S. Regazzoni

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments FUSION 2018 - 21st International Conference on Information Fusion, Cambridge, UK

详情

展开后加载摘要…

URL PDF HTML 收藏