arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 470 信号源:cs.CV, eess.IV, cs.MM

1. 动作与事件理解 470 篇

2408.06437 2024-08-14 cs.CV 57%

HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization

Sakib Reza, Yuexi Zhang, Mohsen Moghaddam, Octavia Camps

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05364 2024-08-13 cs.CV 57%

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, Calvin Murdock

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17792 2024-07-29 cs.CV 57%

Harnessing Temporal Causality for Advanced Temporal Action Detection

Shuming Liu, Lin Sui, Chen-Lin Zhang, Fangzhou Mu, Chen Zhao, Bernard Ghanem

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments 1st in Moment Queries track at the Ego4D Challenge 2024; 1st in Action Recognition, Action Detection, and Audio-Based Interaction Detection tracks at the EPIC-Kitchens Challenge 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07478 2024-07-11 cs.CV 57%

EA-VTR: Event-Aware Video-Text Retrieval

Zongyang Ma, Ziqi Zhang, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Yingmin Luo, Xu Li, Xiaojuan Qi, Ying Shan, Weiming Hu

专题命中 动作与事件理解 :text-to-video(abstract);分类 cs.CV

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06157 2024-07-09 cs.CV cs.AI 57%

Temporal Grounding of Activities using Multimodal Large Language Models

Young Chol Song

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09781 2024-06-17 cs.CV 57%

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou, Jiangong Li

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17984 2024-05-28 cs.CV 57%

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gordon Wetzstein, Leonidas Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, David B. Lindell

专题命中 动作与事件理解 :text-to-video(abstract);分类 cs.CV

Comments CVPR 2024; Project page: https://sherwinbahmani.github.io/4dfy

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13865 2024-05-24 cs.CV 57%

ReVideo: Remake a Video with Motion and Content Control

Chong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang, Ying Shan, Jian Zhang

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07031 2024-05-14 cs.CV 57%

Global Motion Understanding in Large-Scale Video Object Segmentation

Volodymyr Fedynyak, Yaroslav Romanus, Oles Dobosevych, Igor Babin, Roman Riazantsev

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07351 2024-04-15 cs.CV cs.HC cs.LG 57%

A Transformer-Based Model for the Prediction of Human Gaze Behavior on Videos

Suleyman Ozdel, Yao Rong, Berat Mert Albaba, Yen-Ling Kuo, Xi Wang, Enkelejda Kasneci

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments 2024 Symposium on Eye Tracking Research and Applications (ETRA24), Glasgow, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07347 2024-04-15 cs.CV cs.HC cs.LG 57%

Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention

Suleyman Ozdel, Yao Rong, Berat Mert Albaba, Yen-Ling Kuo, Xi Wang, Enkelejda Kasneci

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments 2024 Symposium on Eye Tracking Research and Applications (ETRA24), Glasgow, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05439 2024-04-09 cs.CV 57%

Action-conditioned video data improves predictability

Meenakshi Sarkar, Debasish Ghose

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13185 2024-04-09 cs.CV 57%

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo, Haoji Hu, Zuozhu Liu, Jiang Bian

专题命中 动作与事件理解 :text-to-video(abstract);分类 cs.CV

Comments Project page: https://jianhongbai.github.io/UniEdit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19225 2024-03-29 cs.CV 57%

Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment

Angchi Xu, Wei-Shi Zheng

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07420 2024-03-18 cs.CV 57%

DragAnything: Motion Control for Anything using Entity Representation

Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, Di Zhang

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

Comments The project website is at: https://weijiawu.github.io/draganything_page/ . The code is at: https://github.com/showlab/DragAnything

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02997 2024-02-05 cs.CV 57%

TadML: A fast temporal action detection with Mechanics-MLP

Bowen Deng, Dongchang Liu

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments 8 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12419 2024-01-24 cs.CV 57%

Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)

Shih-Han Chou, Matthew Kowal, Yasmin Niknam, Diana Moyano, Shayaan Mehdi, Richard Pito, Cheng Zhang, Ian Knopke, Sedef Akinli Kocak, Leonid Sigal, Yalda Mohsenzadeh

专题命中 动作与事件理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07207 2023-12-21 cs.CV cs.CL 57%

Beyond Grounding: Extracting Fine-Grained Event Hierarchies Across Modalities

Hammad A. Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh, Long Chen, Yulei Niu, Xudong Lin, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang

专题命中 动作与事件理解 :video-language(abstract);分类 cs.CV

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12250 2023-12-20 cs.CV 57%

ST(OR)2: Spatio-Temporal Object Level Reasoning for Activity Recognition in the Operating Room

Idris Hamoud, Muhammad Abdullah Jamal, Vinkle Srivastav, Didier Mutter, Nicolas Padoy, Omid Mohareri

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03391 2023-12-07 cs.CV 57%

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi, Giovanni Maria Farinella

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12886 2023-12-06 cs.CV 57%

AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance

Zuozhuo Dai, Zhenghao Zhang, Yao Yao, Bingxue Qiu, Siyu Zhu, Long Qin, Weizhi Wang

专题命中 动作与事件理解 :video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00729 2023-11-08 cs.CV cs.AI 57%

ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection

Thinh Phan, Khoa Vo, Duy Le, Gianfranco Doretto, Donald Adjeroh, Ngan Le

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13808 2023-10-11 cs.CV 57%

Streaming Video Temporal Action Segmentation In Real Time

Wujun Wen, Yunheng Li, Zhuben Dong, Lin Feng, Wanxiao Yang, Shenlan Liu

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments accepted by ISKE2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14900 2023-10-10 cs.CV 57%

BIT: Bi-Level Temporal Modeling for Efficient Supervised Action Segmentation

Zijia Lu, Ehsan Elhamifar

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.12317 2023-10-03 cs.CV cs.GR cs.LG 57%

Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis

Chonghyuk Song, Gengshan Yang, Kangle Deng, Jun-Yan Zhu, Deva Ramanan

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments ICCV 2023 camera-ready version. Project page with code, models, and data: https://andrewsonga.github.io/totalrecon

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15086 2023-09-27 cs.CV 57%

Video-adverb retrieval with compositional adverb-action embeddings

Thomas Hummel, Otniel-Bogdan Mercea, A. Sophia Koepke, Zeynep Akata

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments BMVC 2023 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11358 2023-09-26 cs.CV cs.AI cs.LG 57%

How Much Temporal Long-Term Context is Needed for Action Segmentation?

Emad Bahrami, Gianpiero Francesca, Juergen Gall

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15701 2023-09-14 cs.CV 57%

Action Sensitivity Learning for Temporal Action Localization

Jiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng, Jiang Yang, Yi Yang

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00696 2023-09-06 cs.CV 57%

AAN: Attributes-Aware Network for Temporal Action Detection

Rui Dai, Srijan Das, Michael S. Ryoo, Francois Bremond

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09179 2023-08-29 cs.CV cs.AI 57%

Pretrained Language Models as Visual Planners for Human Assistance

Dhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino, Unnat Jain, Ruta Desai

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏