arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2107.05250 2021-11-29 cs.CV 57%

Modeling Explicit Concerning States for Reinforcement Learning in Visual Dialogue

Zipeng Xu, Fandong Meng, Xiaojie Wang, Duo Zheng, Chenxu Lv, Jie Zhou

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments In BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10146 2021-11-22 cs.CV 57%

DVCFlow: Modeling Information Flow Towards Human-like Video Captioning

Xu Yan, Zhengcong Fei, Shuhui Wang, Qingming Huang, Qi Tian

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.09887 2021-11-19 cs.CV cs.LG 57%

PyTorchVideo: A Deep Learning Library for Video Understanding

Haoqi Fan, Tullie Murrell, Heng Wang, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Bo Xiong, Nikhila Ravi, Meng Li, Haichuan Yang, Jitendra Malik, Ross Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, Christoph Feichtenhofer

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.02362 2021-11-17 cs.AI 57%

Video Sentiment Analysis with Bimodal Information-augmented Multi-Head Attention

Ting Wu, Junjie Peng, Wenqiang Zhang, Huiran Zhang, Chuanshuai Ma, Yansong Huang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 12 pages, 4 figures, content and journal information updated

Journal ref Knowledge Based Systems 235 (2022) 107676

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07344 2021-11-16 cs.CV cs.LG 57%

Towards Privacy-Preserving Affect Recognition: A Two-Level Deep Learning Architecture

Jimiama M. Mase, Natalie Leesakul, Fan Yang, Grazziela P. Figueredo, Mercedes Torres Torres

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04052 2021-11-09 cs.CL 57%

How does a Pre-Trained Transformer Integrate Contextual Keywords? Application to Humanitarian Computing

Barriere Valentin, Jacquet Guillaume

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Oral ISCRAM2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.03225 2021-11-08 cs.CV 57%

Technical Report: Disentangled Action Parsing Networks for Accurate Part-level Action Parsing

Xuanhan Wang, Xiaojia Chen, Lianli Gao, Lechao Chen, Jingkuan Song

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02521 2021-11-05 cs.CV q-bio.QM 57%

Sequence-to-Sequence Modeling for Action Identification at High Temporal Resolution

Aakash Kaku, Kangning Liu, Avinash Parnandi, Haresh Rengaraj Rajamohan, Kannan Venkataramanan, Anita Venkatesan, Audre Wirtanen, Natasha Pandit, Heidi Schambra, Carlos Fernandez-Granda

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Under review as a conference paper at ICLR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.08037 2021-10-27 cs.CV 57%

CCVS: Context-aware Controllable Video Synthesis

Guillaume Le Moing, Jean Ponce, Cordelia Schmid

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08061 2021-10-25 cs.CV 57%

Invertible Frowns: Video-to-Video Facial Emotion Translation

Ian Magnusson, Aruna Sankaranarayanan, Andrew Lippman

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 9 pages, 2 figures, 4 tables, accepted at ADGD @ ACM Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01758 2021-10-08 cs.CV 57%

Quantified Facial Expressiveness for Affective Behavior Analytics

Md Taufeeq Uddin, Shaun Canavan

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.03774 2021-10-06 cs.CV 57%

Digital Taxonomist: Identifying Plant Species in Community Scientists' Photographs

Riccardo de Lutio, Yihang She, Stefano D'Aronco, Stefania Russo, Philipp Brun, Jan D. Wegner, Konrad Schindler

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication in the ISPRS Journal of Photogrammetry and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.10024 2021-09-22 cs.RO cs.CV cs.LG 57%

Self-Supervised Action-Space Prediction for Automated Driving

Faris Janjoš, Maxim Dolgov, J. Marius Zöllner

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02602 2021-09-22 cs.LG cs.AI cs.IR 57%

Merchant Category Identification Using Credit Card Transactions

Chin-Chia Michael Yeh, Zhongfang Zhuang, Yan Zheng, Liang Wang, Junpeng Wang, Wei Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09052 2021-09-21 cs.CV 57%

Object Tracking by Jointly Exploiting Frame and Event Domain

Jiqing Zhang, Xin Yang, Yingkai Fu, Xiaopeng Wei, Baocai Yin, Bo Dong

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02925 2021-09-08 cs.CV 57%

Learning to Combine the Modalities of Language and Video for Temporal Moment Localization

Jungkyoo Shin, Jinyoung Moon

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01774 2021-09-08 cs.MM 57%

What Matters for Ad-hoc Video Search? A Large-scale Evaluation on TRECVID

Aozhu Chen, Fan Hu, Zihan Wang, Fangming Zhou, Xirong Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.MM

Comments Accepted by ViRal'21@ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.08027 2021-09-01 cs.CV 57%

DanceIt: Music-inspired Dancing Video Synthesis

Xin Guo, Yifan Zhao, Jia Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 14 pages, 20 figures. [J]. IEEE Transactions on Image Processing, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14371 2021-08-31 cs.CV 57%

Tensor Representations for Action Recognition

Piotr Koniusz, Lei Wang, Anoop Cherian

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Published with TPAMI, 2020. arXiv admin note: text overlap with arXiv:1604.00239

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.14937 2021-08-31 cs.CV 57%

Learning Video Representations from Textual Web Supervision

Jonathan C. Stroud, Zhichao Lu, Chen Sun, Jia Deng, Rahul Sukthankar, Cordelia Schmid, David A. Ross

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.10576 2021-08-25 cs.CV 57%

Support-Set Based Cross-Supervision for Video Grounding

Xinpeng Ding, Nannan Wang, Shiwei Zhang, De Cheng, Xiaomeng Li, Ziyuan Huang, Mingqian Tang, Xinbo Gao

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03329 2021-08-10 cs.CV 57%

Feature-Supervised Action Modality Transfer

Fida Mohammad Thoker, Cees G. M. Snoek

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments IEEE International Conference on Pattern Recognition (ICPR), 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.03298 2021-08-10 cs.CV cs.RO 57%

Bifold and Semantic Reasoning for Pedestrian Behavior Prediction

Amir Rasouli, Mohsen Rohani, Jun Luo

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2021. 11 pages; 5 Figures; 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08307 2021-07-09 cs.CV cs.LG 57%

AC-VRNN: Attentive Conditional-VRNN for Multi-Future Trajectory Prediction

Alessia Bertugli, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at Computer Vision and Image Understanding (CVIU)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00210 2021-07-07 cs.CV 57%

Semantics-aware Adaptive Knowledge Distillation for Sensor-to-Vision Action Recognition

Yang Liu, Keze Wang, Guanbin Li, Liang Lin

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments This paper focuses on the sensor-to-vision heterogenous action recognition problem. Code is available at https://github.com/YangLiu9208/SAKDN

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.06600 2021-07-06 cs.LG cs.AI stat.ML 57%

Zeroth-Order Supervised Policy Improvement

Hao Sun, Ziping Xu, Yuhang Song, Meng Fang, Jiechao Xiong, Bo Dai, Bolei Zhou

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10026 2021-07-02 cs.CV 57%

EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2021: Team M3EM Technical Report

Lijin Yang, Yifei Huang, Yusuke Sugano, Yoichi Sato

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.16136 2021-07-01 cs.CV 57%

Weakly Supervised Temporal Adjacent Network for Language Grounding

Yuechen Wang, Jiajun Deng, Wengang Zhou, Houqiang Li

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10393 2021-06-22 cs.CV 57%

Dynamical Deep Generative Latent Modeling of 3D Skeletal Motion

Amirreza Farnoosh, Sarah Ostadabbas

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.00391 2021-06-21 cs.CV 57%

A Novel Graph based Trajectory Predictor with Pseudo Oracle

Biao Yang, Guocheng Yan, Pin Wang, Chingyao Chan, Xiang Song, Yang Chen

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 17 apges, 8 figures

Journal ref already published by TNNLS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏