arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2206.09852 2022-06-22 cs.CV 79%

M&M Mix: A Multimodal Multiview Transformer Ensemble

Xuehan Xiong, Anurag Arnab, Arsha Nagrani, Cordelia Schmid

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Technical report for Epic-Kitchens challenge 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02353 2022-06-09 cs.LG cs.CV 79%

Beyond Just Vision: A Review on Self-Supervised Representation Learning on Multimodal and Temporal Data

Shohreh Deldari, Hao Xue, Aaqib Saeed, Jiayuan He, Daniel V. Smith, Flora D. Salim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 36 pages, 5 figures, 9 tables, Survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01657 2022-05-04 cs.CV 79%

Cross-modal Representation Learning for Zero-shot Action Recognition

Chung-Ching Lin, Kevin Lin, Linjie Li, Lijuan Wang, Zicheng Liu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15086 2022-03-30 cs.CV 79%

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13301 2022-03-29 cs.CV 79%

Multi-modal Multi-label Facial Action Unit Detection with Transformer

Lingfeng Wang, Shisen Wang, Jin Qi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12745 2022-03-29 cs.CV 79%

UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Ye Liu, Siyuan Li, Yang Wu, Chang Wen Chen, Ying Shan, Xiaohu Qie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01122 2022-03-22 cs.CV 79%

M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers

Tsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09613 2022-03-18 cs.CV eess.IV 79%

SEN12MS-CR-TS: A Remote Sensing Data Set for Multi-modal Multi-temporal Cloud Removal

Patrick Ebel, Yajin Xu, Michael Schmitt, Xiaoxiang Zhu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Geoscience and Remote Sensing, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09124 2022-02-21 eess.AS cs.SD 79%

Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion

Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada

专题命中 视频多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 5 pages, 5 figures, to appear in IEEE ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03283 2022-02-08 cs.CV 79%

CZU-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and 10 wearable inertial sensors

Xin Chao, Zhenjie Hou, Yujian Mo

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.12713 2022-01-25 cs.CV 79%

Spatio-Contextual Deep Network Based Multimodal Pedestrian Detection For Autonomous Driving

Kinjal Dasgupta, Arindam Das, Sudip Das, Ujjwal Bhattacharya, Senthil Yogamani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be published at IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08443 2021-12-17 cs.LG cs.AI 79%

Event-Aware Multimodal Mobility Nowcasting

Zhaonan Wang, Renhe Jiang, Hao Xue, Flora D. Salim, Xuan Song, Ryosuke Shibasaki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07558 2021-12-15 cs.CV eess.IV 79%

Multi-Modal Temporal Attention Models for Crop Mapping from Satellite Time Series

Vivien Sainte Fare Garnot, Loic Landrieu, Nesrine Chehata

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05379 2021-12-13 cs.CV cs.CR cs.LG 79%

Cross-Modal Transferable Adversarial Attacks from Images to Videos

Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01677 2021-11-03 cs.CV cs.LG 79%

Top1 Solution of QQ Browser 2021 Ai Algorithm Competition Track 1 : Multimodal Video Similarity

Zhuoran Ma, Majing Lou, Xuan Ouyang

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00865 2021-11-02 cs.CV eess.IV 79%

MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition

Jinming Zhao, Ruichen Li, Qin Jin, Xinchao Wang, Haizhou Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 4 papges, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05624 2021-11-01 cs.CV 79%

End-to-end Multi-modal Video Temporal Grounding

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in NeurIPS 2021. Project page at https://github.com/wenz116/DRFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08270 2021-10-20 cs.LG cs.CL 79%

From Multimodal to Unimodal Attention in Transformers using Knowledge Distillation

Dhruv Agarwal, Tanay Agrawal, Laura M. Ferrari, François Bremond

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Preprint. Final paper accepted at the 17th IEEE International Conference on Advanced Video and Signal-based Surveillance, AVSS 2021, Virtual, November 16-19, 2021. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08814 2021-10-19 cs.CV 79%

TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding

Zhengwei Wang, Qi She, Aljosa Smolic

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments To appear in BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12671 2021-10-18 cs.CV 79%

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, Michael Picheny, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be presented at ICCV 2021

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 8012-8021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06058 2021-10-13 cs.CV 79%

Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

Zongmeng Zhang, Xianjing Han, Xuemeng Song, Yan Yan, Liqiang Nie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing

Journal ref in IEEE Transactions on Image Processing, vol. 30, pp. 8265-8277, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13666 2021-09-29 cs.CV cs.RO 79%

Fail-Safe Human Detection for Drones Using a Multi-Modal Curriculum Learning Approach

Ali Safa, Tim Verbelen, Ilja Ocket, André Bourdoux, Francky Catthoor, Georges G. E. Gielen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04735 2021-09-13 cs.CV 79%

Temporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

Min Peng, Chongyang Wang, Yuan Gao, Yu Shi, Xiang-Dong Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to AAAI'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03619 2021-08-10 cs.CV 79%

Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action Detection

Rui Dai, Srijan Das, Francois Bremond

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12589 2021-07-28 cs.CV 79%

Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization

Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, Wei-Shi Zheng

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ACM International Conference on Multimedia, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14435 2021-07-22 cs.CV 79%

DanHAR: Dual Attention Network For Multimodal Human Activity Recognition Using Wearable Sensors

Wenbin Gao, Lei Zhang, Qi Teng, Jun He, Hao Wu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09504 2021-07-21 cs.CV 79%

Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos

Olga Zatsarynna, Yazan Abu Farha, Juergen Gall

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR Precognition Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04187 2021-07-16 cs.CV 79%

A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition

Yue Jin, Tianqing Zheng, Chao Gao, Guoqiang Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 5 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03009 2021-07-13 cs.CV 79%

Multi-modal Affect Analysis using standardized data within subjects in the Wild

Sachihiro Youoku, Takahisa Yamamoto, Junya Saito, Akiyoshi Uchida, Xiaoyu Mi, Ziqiang Shi, Liu Liu, Zhongling Liu, Osafumi Nakayama, Kentaro Murase

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09412 2021-07-07 cs.CV 79%

Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition

Zitong Yu, Benjia Zhou, Jun Wan, Pichao Wang, Haoyu Chen, Xin Liu, Stan Z. Li, Guoying Zhao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Submitted to IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏