arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6759 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1375 篇

2404.03398 2024-04-05 cs.CV 57%

Scaling Up Video Summarization Pretraining with Large Language Models

Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron, Hanieh Deilamsalehy, Trung Bui, Zhaowen Wang, Franck Dernoncourt, Joon Son Chung

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01297 2024-04-02 cs.CV 57%

Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, Cordelia Schmid

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments CVPR 2024. Code is available at https://github.com/google-research/scenic/tree/main/scenic/projects/streaming_dvc

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00901 2024-04-02 cs.CV 57%

Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, Fahad Shahbaz Khan

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14326 2024-03-28 cs.MM 57%

Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation

Mingxuan Yan, Yi Wang, Xuedou Xiao, Zhiqing Luo, Jianhua He, Wei Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09506 2024-03-26 cs.CV cs.AI cs.LG 57%

Don't Judge by the Look: Towards Motion Coherent Video Representation

Yitian Zhang, Yue Bai, Huan Wang, Yizhou Wang, Yun Fu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICLR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04940 2024-03-11 cs.CV cs.AI cs.LG q-bio.NC 57%

A spatiotemporal style transfer algorithm for dynamic visual stimulus generation

Antonino Greco, Markus Siegel

专题命中 视频理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19479 2024-03-01 cs.CV 57%

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, Sergey Tulyakov

专题命中 视频理解 :video generation(abstract);分类 cs.CV

Comments CVPR 2024. Project Page: https://snap-research.github.io/Panda-70M

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02574 2024-02-07 cs.CV cs.LG 57%

Spatio-temporal Prompting Network for Robust Video Feature Extraction

Guanxiong Sun, Chi Wang, Zhaoyu Zhang, Jiankang Deng, Stefanos Zafeiriou, Yang Hua

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Journal ref 2023 International Conference on Computer Vision (ICCV) 13541-13551

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10254 2024-01-22 cs.CV cs.LG 57%

Beyond the Frame: Single and mutilple video summarization method with user-defined length

Vahid Ahmadi Kalkhorani, Qingquan Zhang, Guanqun Song, Ting Zhu

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10702 2024-01-22 cs.CV 57%

ClawCraneNet: Leveraging Object-level Relation for Text-based Video Segmentation

Chen Liang, Yu Wu, Yawei Luo, Yi Yang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Extended version published in https://ieeexplore.ieee.org/abstract/document/10083244

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13008 2023-12-21 cs.CV cs.AI cs.LG 57%

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision

Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments AAAI 2024 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07378 2023-12-13 cs.CV 57%

X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-modal Knowledge Transfer

Linglin Jing, Ying Xue, Xu Yan, Chaoda Zheng, Dong Wang, Ruimao Zhang, Zhigang Wang, Hui Fang, Bin Zhao, Zhen Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04821 2023-12-08 cs.CV 57%

PromptonomyViT: Multi-Task Prompt Learning Improves Video Transformers using Synthetic Scene Data

Roei Herzig, Ofir Abramovich, Elad Ben-Avraham, Assaf Arbelle, Leonid Karlinsky, Ariel Shamir, Trevor Darrell, Amir Globerson

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17043 2023-11-29 cs.CV cs.CL 57%

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Yanwei Li, Chengyao Wang, Jiaya Jia

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Code is available at https://github.com/dvlab-research/LLaMA-VID

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11365 2023-11-13 cs.CV 57%

EgoEnv: Human-centric environment representations from egocentric video

Tushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis, Kristen Grauman

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Published in NeurIPS 2023 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02559 2023-11-07 cs.CV 57%

Rotation Invariant Transformer for Recognizing Object in UAVs

Shuoyi Chen, Mang Ye, Bo Du

专题命中 视频理解 :video reasoning(abstract);分类 cs.CV

Comments ACM MM2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08763 2023-10-31 cs.CV 57%

Video-Mined Task Graphs for Keystep Recognition in Instructional Videos

Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen Grauman

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04991 2023-10-12 cs.CV 57%

Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling

Haogeng Liu, Qihang Fan, Tingkai Liu, Linjie Yang, Yunzhe Tao, Huaibo Huang, Ran He, Hongxia Yang

专题命中 视频理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10783 2023-09-20 cs.CV cs.AI cs.CL 57%

Language as the Medium: Multimodal Video Classification through text only

Laura Hanu, Anita L. Verő, James Thewlis

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted at "What is Next in Multimodal Foundation Models?" (MMFM) workshop at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14965 2023-08-30 cs.CV 57%

CEFHRI: A Communication Efficient Federated Learning Framework for Recognizing Industrial Human-Robot Interaction

Umar Khalid, Hasan Iqbal, Saeed Vahidian, Jing Hua, Chen Chen

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted in IROS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14274 2023-08-29 cs.MM 57%

Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks

Hongye Liu, Xianhai Xie, Yang Gao, Size Li, Zhou YU

专题命中 视频理解 :video understanding(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09245 2023-08-21 cs.CV cs.AI 57%

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan, Longguang Wang, Yulan Guo, Qiong Liu, Hao Wen, Xi Zhou

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07918 2023-08-16 cs.CV 57%

Helping Hands: An Object-Aware Ego-Centric Video Recognition Model

Chuhan Zhang, Ankush Gupta, Andrew Zisserman

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08265 2023-08-08 cs.CV cs.LG 57%

Vehicle Detection and Classification without Residual Calculation: Accelerating HEVC Image Decoding with Random Perturbation Injection

Muhammet Sebul Beratoğlu, Behçet Uğur Töreyin

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 10 pages 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03740 2023-07-19 cs.CV 57%

Convolutional Hierarchical Attention Network for Query-Focused Video Summarization

Shuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan, Min Yang

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Accepted by AAAI 2020 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07483 2023-07-19 cs.CV 57%

Multimodal Distillation for Egocentric Action Recognition

Gorjan Radevski, Dusan Grujicic, Marie-Francine Moens, Matthew Blaschko, Tinne Tuytelaars

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted at ICCV 2023; Codebase released at https://github.com/gorjanradevski/multimodal-distillation

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12005 2023-07-06 cs.CL cs.CV 57%

mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si

专题命中 视频理解 :video-language(abstract);分类 cs.CV

Journal ref EMNLP2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13118 2023-06-26 cs.AI cs.CV cs.IR 57%

An overview on the evaluated video retrieval tasks at TRECVID 2022

George Awad, Keith Curtis, Asad Butt, Jonathan Fiscus, Afzal Godil, Yooyoung Lee, Andrew Delgado, Eliot Godard, Lukas Diduch, Jeffrey Liu, Yvette Graham, Georges Quenot

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2104.13473, arXiv:2009.09984

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13292 2023-05-24 cs.CV 57%

VideoLLM: Modeling Video Sequence with Large Language Models

Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, Limin Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06378 2023-05-18 cs.CV 57%

Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos

Teng Wang, Jinrui Zhang, Feng Zheng, Wenhao Jiang, Ran Cheng, Ping Luo

专题命中 视频理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏