arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6737 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1369 篇

2406.16301 2024-06-25 cs.CV cs.AI cs.MM 62%

UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos

Yuting Mei, Linli Yao, Qin Jin

专题命中 视频理解 :long video(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM International Conference on Multimedia Retrieval (ICMR'24)

Journal ref Proceedings of the 2024 International Conference on Multimedia Retrieval, May 2024, Pages 1034-1042

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04192 2024-04-02 cs.CV cs.AI cs.CL cs.MM 62%

Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models

Wei Han, Hui Chen, Min-Yen Kan, Soujanya Poria

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures, accepted to Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12968 2023-11-23 cs.CV cs.MM 62%

Improved Actor Relation Graph based Group Activity Recognition

Zijian Kuang, Xinran Tie

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Journal ref ICSM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00282 2023-11-13 cs.CV cs.IR cs.MM 62%

(Un)likelihood Training for Interpretable Embedding

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Zhijian Hou

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments accepted in ACM Transactions on Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10227 2023-09-20 eess.IV cs.CV 62%

Learning Dynamic MRI Reconstruction with Convolutional Network Assisted Reconstruction Swin Transformer

Di Xu, Hengjie Liu, Dan Ruan, Ke Sheng

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments MICCAI 2023 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13004 2023-08-28 cs.CV cs.AI cs.MM 62%

Spherical Vision Transformer for 360-degree Video Saliency Prediction

Mert Cokelek, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 12 pages, 4 figures, accepted to BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09322 2023-08-21 cs.CV cs.AI cs.MM 62%

Audio-Visual Glance Network for Efficient Video Recognition

Muhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick Kim

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03063 2023-08-08 cs.CV cs.MM 62%

M$^3$Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition

Hao Tang, Jun Liu, Shuanglin Yan, Rui Yan, Zechao Li, Jinhui Tang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13176 2023-06-26 cs.CV cs.LG eess.IV 62%

Key Frame Extraction with Attention Based Deep Neural Networks

Samed Arslan, Senem Tanberk

专题命中 视频理解 :long video(abstract);分类 cs.CV、eess.IV

Comments in Turkish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12037 2023-03-23 cs.CV cs.AI cs.LG cs.MM 62%

Causal Reasoning Meets Visual Representation Learning: A Prospective Study

Yang Liu, Yushen Wei, Hong Yan, Guanbin Li, Liang Lin

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 35 pages, 14 figures. This work has been accepted by Machine Intelligence Research. The arxiv version is kept updating by adding more novel methods, datasets and insights. The official video interpretation of this paper can be referred at https://youtu.be/2lfNaTkcTHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04979 2023-03-17 cs.CV cs.LG cs.MM 62%

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, Jiahui Yu

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV、cs.MM

Comments Tech report. arXiv v3: update text

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04752 2022-11-15 eess.IV cs.CV 62%

Human Gaze Guided Attention for Surgical Activity Recognition

Abdishakour Awale, Duygu Sarikaya

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04212 2022-07-19 cs.CV cs.LG eess.IV 62%

AutoVideo: An Automated Video Action Recognition System

Daochen Zha, Zaid Pervaiz Bhat, Yi-Wei Chen, Yicheng Wang, Sirui Ding, Jiaben Chen, Kwei-Herng Lai, Mohammad Qazim Bhat, Anmoll Kumar Jain, Alfredo Costilla Reyes, Na Zou, Xia Hu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments Accepted by IJCAI https://github.com/datamllab/autovideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.06931 2022-06-15 cs.CV cs.AI cs.MM 62%

Stand-Alone Inter-Frame Attention in Video Models

Fuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao, Jiebo Luo, Tao Mei

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments CVPR 2022; Code is publicly available at: https://github.com/FuchenUSTC/SIFA

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03838 2022-03-09 cs.CV cs.AI cs.LG cs.MM 62%

Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding

Shentong Mo, Daizong Liu, Wei Hu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00629 2021-11-02 cs.MM cs.CV cs.IR cs.LG 62%

Distantly Supervised Semantic Text Detection and Recognition for Broadcast Sports Videos Understanding

Avijit Shah, Topojoy Biswas, Sathish Ramadoss, Deven Santosh Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 9 pages, 7 figures and 6 tables. To be published in the proceedings of ACM Multimedia 21, Industrial Track, held from October 20-24 in China

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06637 2021-09-15 cs.CV cs.MM 62%

Multi-modal Representation Learning for Video Advertisement Content Structuring

Daya Guo, Zhaoyang Zeng

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02432 2021-08-06 cs.CV cs.MM 62%

Token Shift Transformer for Video Classification

Hao Zhang, Yanbin Hao, Chong-Wah Ngo

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2021, 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06567 2020-12-14 cs.CV cs.MM 62%

A Comprehensive Study of Deep Video Action Recognition

Yi Zhu, Xinyu Li, Chunhui Liu, Mohammadreza Zolfaghari, Yuanjun Xiong, Chongruo Wu, Zhi Zhang, Joseph Tighe, R. Manmatha, Mu Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Technical report. Code and model zoo can be found at https://cv.gluon.ai/model_zoo/action_recognition.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10260 2020-06-19 cs.CV cs.MM 62%

Video Moment Localization using Object Evidence and Reverse Captioning

Madhawa Vidanapathirana, Supriya Pandhre, Sonia Raychaudhuri, Anjali Khurana

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 7 pages. 6 figures. For source code, refer https://github.com/madhawav/MML

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07613 2020-01-22 cs.LG cs.CV eess.IV stat.ML 62%

Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video Data

Kyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments 8 pages, 3 figures, Association for the Advancement of Artificial Intelligence (AAAI2020). arXiv admin note: substantial text overlap with arXiv:1907.01709

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02793 2019-10-08 cs.CV cs.LG eess.IV 62%

ViP: Video Platform for PyTorch

Madan Ravi Ganesh, Eric Hofesmann, Nathan Louis, Jason Corso

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.12522 2018-10-31 cs.CV cs.AI cs.MM 62%

Random Temporal Skipping for Multirate Video Analysis

Yi Zhu, Shawn Newsam

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted at ACCV 2018. Camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.06858 2017-08-24 cs.MM cs.CV cs.HC 62%

ElasticPlay: Interactive Video Summarization with Dynamic Time Budgets

Haojian Jin, Yale Song, Koji Yatani

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2017 preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17058 2023-11-29 cs.CV cs.AI 61%

Panoptic Video Scene Graph Generation

Jingkang Yang, Wenxuan Peng, Xiangtai Li, Zujin Guo, Liangyu Chen, Bo Li, Zheng Ma, Kaiyang Zhou, Wayne Zhang, Chen Change Loy, Ziwei Liu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV;long video(comments)

Comments Accepted to CVPR 2023. Project Page: https://jingkang50.github.io/PVSG/. Codebase: https://github.com/LilyDaytoy/OpenPVSG. We provide 400 long videos with frame-level panoptic segmentation, scene graph, dense captions, and QA annotations

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13380 2023-06-26 cs.CV 61%

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li, Zechuan Li, Jingwen Wang, Wei Miao, Wei Sun, Chen Chen

专题命中 视频理解 :video-language(abstract);分类 cs.CV;video understanding(comments)

Comments Winner of CVPR2023 Long-form Video Understanding and Generation Challenge (Track 3)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06866 2019-11-19 cs.CV 61%

Multi-attention Networks for Temporal Localization of Video-level Labels

Lijun Zhang, Srinath Nizampatnam, Ahana Gangopadhyay, Marcos V. Conde

专题命中 视频理解 :video understanding(abstract,comments);分类 cs.CV

Comments 7 pages, 3 figures; This work was presented at the 3rd Workshop on YouTube-8M Large-Scale Video Understanding, at the International Conference on Computer Vision (ICCV 2019) in Seoul, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.00207 2018-10-02 cs.CV 61%

Non-local NetVLAD Encoding for Video Classification

Yongyi Tang, Xing Zhang, Jingwen Wang, Shaoxiang Chen, Lin Ma, Yu-Gang Jiang

专题命中 视频理解 :video understanding(abstract,comments);分类 cs.CV

Comments ECCV2018 workshop on YouTube-8M Large-Scale Video Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.09274 2017-07-14 cs.CV 61%

The YouTube-8M Kaggle Competition: Challenges and Methods

Haosheng Zou, Kun Xu, Jialian Li, Jun Zhu

专题命中 视频理解 :video understanding(abstract,comments);分类 cs.CV

Comments accepted to CVPR'17 Workshop on YouTube-8M Large-Scale Video Understanding (oral presentation); code is at https://github.com/taufikxu/youtube on branches kunxu and zhs

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14262 2026-08-17 cs.CV 新提交 57%

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

手术内窥镜视频的时间视觉语言模型的鲁棒性研究

Darakshan Rashid, Raza Imam, Ufaq Khan, Muhammad Bilal, Shazad Ashraf, Dwarikanath Mahapatra, Mohammad Yaqub, Muhammad Haris Khan, Imran Razzak, Brejesh Lall, Lena Maier-Hein, Yutong Xie

机构 * Indian Institute of Technology Delhi(印度理工学院德里分校) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Birmingham City University(伯明翰城市大学) University Hospitals Birmingham(伯明翰大学医院) Khalifa University(哈利法大学) German Cancer Research Center (DKFZ)(德国癌症研究中心)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 该研究针对手术内窥镜视频的时间视觉语言模型,构建Endo-C6基准测试其鲁棒性,提出RobustEndoCLIP,可提升模型在内窥镜损坏下的性能与鲁棒性。

Comments Accepted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏