arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6759 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1375 篇

2106.11297 2022-04-05 cs.CV cs.LG 57%

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Michael S. Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, Anelia Angelova

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments This is the full version of the paper, extending its conference paper at NeurIPS 2021. Version 1.1 of the code is released

Journal ref NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09212 2022-04-01 cs.CV cs.AI 57%

Long-Short Temporal Contrastive Learning of Video Transformers

Jue Wang, Gedas Bertasius, Du Tran, Lorenzo Torresani

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07974 2022-03-31 cs.CV 57%

TCLR: Temporal Contrastive Learning for Video Representation

Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to Computer Vision and Image Understanding (CVIU) Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15205 2022-03-30 cs.CV cs.CR cs.LG 57%

SPAct: Self-supervised Privacy Preservation for Action Recognition

Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15086 2022-03-30 cs.CV 57%

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu

专题命中 视频理解 :video reasoning(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04850 2022-03-18 cs.CV 57%

Bridging Video-text Retrieval with Multiple Choice Questions

Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, Ping Luo

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02573 2022-03-08 cs.CV cs.AI cs.LG 57%

Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning

Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, Sergey Tulyakov

专题命中 视频理解 :video generation(abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10828 2022-02-23 cs.CV 57%

Exploiting long-term temporal dynamics for video captioning

Yuyu Guo, Jingqiu Zhang, Lianli Gao

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09979 2022-02-22 cs.CL cs.CV 57%

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida, Akihiko Takashima

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted at DSTC10 Workshop at AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12086 2022-02-16 cs.CV 57%

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi

专题命中 视频理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09153 2022-01-25 cs.CV cs.AI 57%

An Integrated Approach for Video Captioning and Applications

Soheyla Amirian, Thiab R. Taha, Khaled Rasheed, Hamid R. Arabnia

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments The 2021 World Congress in Computer Science, Computer Engineering, and Applied Computing (CSCE'21), IEEE, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02359 2022-01-11 cs.CL cs.CV 57%

O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning

Fenglin Liu, Xuancheng Ren, Xian Wu, Bang Yang, Shen Ge, Yuexian Zou, Xu Sun

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by Findings of ACL 2021 (The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08913 2021-12-21 cs.CV 57%

Contrastive Spatio-Temporal Pretext Learning for Self-supervised Video Representation

Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing-Yin Yu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by AAAI 2022, Preprint version with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06423 2021-12-17 cs.CV 57%

Zero-Shot Action Recognition in Videos: A Survey

Valter Estevam, Helio Pedrini, David Menotti

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03803 2021-12-09 cs.CV 57%

Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning

Manlin Zhang, Jinpeng Wang, Andy J. Ma

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments AAAI2022. v2: Add supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01038 2021-12-03 cs.CV 57%

Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips

Lijin Yang, Yifei Huang, Yusuke Sugano, Yoichi Sato

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14785 2021-12-02 cs.CV cs.CL 57%

A Comprehensive Review of the Video-to-Text Problem

Jesus Perez-Martin, Benjamin Bustos, Silvio Jamil F. Guimarães, Ivan Sipiran, Jorge Pérez, Grethel Coello Said

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

Comments 66 pages, 6 figures. Accepted by Artificial Intelligence Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13992 2021-10-28 cs.CV cs.LG 57%

Leveraging Local Temporal Information for Multimodal Scene Classification

Saurabh Sahu, Palash Goyal

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.13384 2021-10-27 cs.CV 57%

ViDA-MAN: Visual Dialog with Digital Humans

Tong Shen, Jiawei Zuo, Fan Shi, Jin Zhang, Liqin Jiang, Meng Chen, Zhengchen Zhang, Wei Zhang, Xiaodong He, Tao Mei

专题命中 视频理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06615 2021-10-14 cs.CV 57%

CLIP4Caption: CLIP for Video Caption

Mingkang Tang, Zhanyu Wang, Zhenhua Liu, Fengyun Rao, Dian Li, Xiu Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.14503 2021-10-11 cs.CV 57%

End-to-End Video Instance Segmentation with Transformers

Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, Huaxia Xia

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR2021 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01015 2021-10-05 cs.CV cs.AI cs.LG 57%

Spatio-Temporal Video Representation Learning for AI Based Video Playback Style Prediction

Rishubh Parihar, Gaurav Ramola, Ranajit Saha, Ravi Kini, Aniket Rege, Sudha Velusamy

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 10 pages, 5 figures, 4 tables, ICCV Workshops 2021 - SRVU

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.11593 2021-09-27 cs.CV 57%

Long Short View Feature Decomposition via Contrastive Video Representation Learning

Nadine Behrmann, Mohsen Fayyaz, Juergen Gall, Mehdi Noroozi

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments ICCV 2021 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00222 2021-09-02 cs.CV 57%

Spatio-Temporal Perturbations for Video Attribution

Zhenqiang Li, Weimin Wang, Zuoyue Li, Yifei Huang, Yoichi Sato

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02833 2021-08-20 cs.CV 57%

Elaborative Rehearsal for Zero-shot Action Recognition

Shizhe Chen, Dong Huang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02183 2021-08-18 cs.CV 57%

Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization

Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, Weiyao Lin

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments ICCV 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.02963 2021-07-28 cs.CV 57%

Disentangle Your Dense Object Detector

Zehui Chen, Chenhongyi Yang, Qiaofei Li, Feng Zhao, Zheng-Jun Zha, Feng Wu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments ACM MM2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06961 2021-07-01 cs.CV 57%

Tiny Video Networks

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14447 2021-06-29 cs.CV cs.AI cs.LG 57%

Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection

Xin Zhou, Le Kang, Zhiyu Cheng, Bo He, Jingyu Xin

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Tech Report. Authors Xin Zhou, Le Kang, and Zhiyu Cheng made equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11250 2021-06-22 cs.CV cs.AI cs.LG 57%

VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Hao Tan, Jie Lei, Thomas Wolf, Mohit Bansal

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Under review, 23 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏