arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6781 信号源:cs.CV, eess.IV, cs.MM

1. 动作与事件理解 473 篇

2605.23355 2026-05-25 cs.CV cs.LG cs.MM 62%

Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization

解耦时空适配器用于细粒度羽毛球动作定位

Tianyu Wang, Junjie Wu, Jingquan Gao, Shishuo Li

机构 * School of Economics and Management, Beihang University(北京航空航天大学经济管理学院) Key Laboratory of Data Intelligence and Management, Beihang University, Ministry of Industry and Information Technology(信息产业部北京航空航天大学数据智能与管理重点实验室)

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、cs.MM

AI总结 针对细粒度羽毛球动作定位中复杂的时空动态,提出解耦时空适配器(DSTA),通过三个并行分支分别建模时间动态、垂直和水平空间变化,在Fine-Badminton数据集和ShuttleSet基准上实现最优性能且计算开销极小。

Comments 11 pages, 11figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15307 2026-05-18 cs.GR cs.CV cs.MM cs.SD 62%

Sound Sparks Motion: Audio and Text Tuning for Video Editing

声音激发动作:用于视频编辑的音频和文本微调

AmirHossein Naghi Razlighi, Aryan Mikaeili, Ali Mahdavi-Amiri, Daniel Cohen-Or, Yiorgos Chrysanthou

机构 * University of Cyprus(塞浦路斯大学) Simon Fraser University(西蒙弗雷泽大学) Tel Aviv University(特拉维夫大学) CYENS Center Of Excellence(CYENS卓越中心)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV、cs.MM

AI总结 本文提出Sound Sparks Motion框架,通过测试时调整音频视觉生成模型的多模态条件信号,实现视频动作编辑,无需训练,通过音频潜在和文本条件残差扰动促进动作修改,同时利用视觉语言模型反馈提升编辑效果。

Comments Project Page: https://amirhossein-razlighi.github.io/Sound_Sparks_Motion

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20169 2026-03-23 cs.CV cs.MM 62%

EgoForge: Goal-Directed Egocentric World Simulator

EgoForge:面向目标的自体世界模拟器

Yifan Shen, Jiateng Liu, Xinzhuo Li, Yuanzhe Liu, Bingxuan Li, Houze Yang, Wenqi Jia, Yijiang Li, Tianjiao Yu, James Matthew Rehg, Xu Cao, Ismini Lourentzou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California San Diego(加州大学圣地亚哥分校)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV、cs.MM

AI总结 EgoForge通过最小静态输入生成连贯的第一人称视频,结合VideoDiffusionNFT提升意图对齐和时间一致性,实现在动态环境模拟中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12060 2025-12-16 cs.CV cs.LG cs.MM 62%

CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos

CreativeVR: 一种基于扩散先验的结构与运动修复方法用于生成视频和真实视频

Tejas Panambur, Ishan Rajendrakumar Dave, Chongjian Ge, Ersin Yumer, Xue Bai

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Adobe(Adobe公司)

专题命中 动作与事件理解 :text-to-video(abstract);分类 cs.CV、cs.MM

AI总结 CreativeVR通过扩散先验引导方法,有效修复AI生成和真实视频中的结构与运动伪影,实现高质量视频修复。

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03179 2025-11-20 cs.CV cs.MM cs.SD eess.AS 62%

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan, Yanfu Zhang, Weili Guan, Feng Zheng, Ling shao

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) William and Mary(威廉与玛丽学院) Harbin Institute of Technology(哈尔滨工业大学) UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Published on IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06855 2025-10-09 cs.CV eess.IV 62%

Online Generic Event Boundary Detection

Hyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son, Jonghyun Choi

机构 * GIST(韩国科学技术院) Seoul National University(首尔国立大学) POSTECH

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06884 2025-04-10 cs.MM cs.AI cs.CV 62%

Audio-visual Event Localization on Portrait Mode Short Videos

Wuyang Liu, Yi Chai, Yongpeng Yan, Yanzhen Ren

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15279 2024-10-22 cs.CV cs.AI cs.MM 62%

ContextDet: Temporal Action Detection with Adaptive Context Aggregation

Ning Wang, Yun Xiao, Xiaopeng Peng, Xiaojun Chang, Xuanhong Wang, Dingyi Fang

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04629 2024-06-10 cs.CV cs.GR cs.MM 62%

STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting

Zenghao Chai, Chen Tang, Yongkang Wong, Mohan Kankanhalli

专题命中 动作与事件理解 :text-to-video(abstract);分类 cs.CV、cs.MM

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14053 2023-03-29 cs.CV cs.AI cs.MM 62%

Re^2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization

Chen Zhao, Shuming Liu, Karttikeya Mangalam, Bernard Ghanem

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13717 2022-08-31 cs.CV cs.AI cs.LG cs.MM cs.SD 62%

StableFace: Analyzing and Improving Motion Stability for Talking Face Generation

Jun Ling, Xu Tan, Liyang Chen, Runnan Li, Yuchao Zhang, Sheng Zhao, Li Song

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV、cs.MM

Comments 12 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.11460 2022-04-20 cs.CV cs.LG eess.IV 62%

Perceptron Synthesis Network: Rethinking the Action Scale Variances in Videos

Yuan Tian, Guangtao Zhai, Zhiyong Gao

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.02463 2021-07-20 cs.CV cs.LG eess.IV 62%

Spatio-Temporal Event Segmentation and Localization for Wildlife Extended Videos

Ramy Mounir, Roman Gula, Jörn Theuerkauf, Sudeep Sarkar

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07976 2021-04-22 cs.CV cs.LG eess.IV 62%

Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization

Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, Hongsheng Li

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments Accepted in CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02401 2020-07-20 cs.CV cs.LG eess.IV 62%

Generating Videos of Zero-Shot Compositions of Actions and Objects

Megha Nawhal, Mengyao Zhai, Andreas Lehrmann, Leonid Sigal, Greg Mori

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV、eess.IV

Comments Accepted at ECCV'20; Project Page: https://www.sfu.ca/~mnawhal/projects/zs_hoi_generation.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03538 2020-04-24 cs.CV cs.LG eess.IV q-bio.PE 62%

Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection

Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, Jonathan Huang

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.01326 2020-03-31 cs.CV cs.LG eess.IV 62%

A Context-Aware Loss Function for Action Spotting in Soccer Videos

Anthony Cioppa, Adrien Deliège, Silvio Giancola, Bernard Ghanem, Marc Van Droogenbroeck, Rikke Gade, Thomas B. Moeslund

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments Accepted for CVPR2020 main conference. This document contains 8 pages + references + supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06699 2019-10-16 cs.CV cs.LG cs.MM 62%

Generating Human Action Videos by Coupling 3D Game Engines and Probabilistic Graphical Models

César Roberto de Souza, Adrien Gaidon, Yohann Cabon, Naila Murray, Antonio Manuel López

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV、cs.MM

Comments Pre-print of the article accepted for publication in the Special Issue on Generating Realistic Visual Data of Human Behavior of the International Journal of Computer Vision (IJCV). arXiv admin note: substantial text overlap with arXiv:1612.00881

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09178 2022-06-22 cs.CV cs.AI 61%

REVECA -- Rich Encoder-decoder framework for Video Event CAptioner

Jaehyuk Heo, YongGi Jeong, Sunwoo Kim, Jaehee Kim, Pilsung Kang

专题命中 动作与事件理解 :video understanding(abstract,comments);分类 cs.CV

Comments The IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR). LOng-form VidEo Understanding (LOVEU) workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16354 2026-08-18 cs.AI cs.CV 新提交 57%

DriveCache: Action-Aware Caching for Driving World Model Inference

DriveCache:面向驾驶世界模型推理的动作感知缓存

Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 针对扩散驾驶生成器吞吐量受限的问题,提出动作感知的DriveCache控制器,利用规划运动与动态规划优化缓存,提升保真度-效率权衡,代码将公开。

Comments 9 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15058 2026-08-18 cs.CV 新提交 57%

MEDR: Query-Independent Frame Selection via Multi-Signal Event Modeling and Dynamic Rescoring

MEDR:基于多信号事件建模与动态重评分的查询无关帧选择

Xinlei Pu, Weijie Shi, Wen Yang, Yi Cao, Hao Chen, Yuanjun Liu, Wenwei Ding, Jia Zhu, Jiajie Xu

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

AI总结 该研究提出无需训练的MEDR帧选择方法,通过多信号事件建模与动态重评分实现查询无关,在Video-MME等基准上提升了多模态大语言模型的长视频处理准确率,且帧集可跨问题复用。

Comments 9 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14022 2026-08-17 cs.CV cs.AI 新提交 57%

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

ForgeWM:面向少步骤动作条件视频世界模型的渐进式因果训练

Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam

机构 * CUHK(香港中文大学) Tencent PCG(腾讯平台与内容事业群) FDU(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) HKUST(香港科技大学)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 ForgeWM 是一种渐进式框架,通过多阶段技术将双向动作条件视频生成器转化为少步骤世界模型,在 Minecraft 和 FPS 游戏任务中实现了更优的可控少步骤视频生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27101 2026-08-13 cs.CV cs.CL 版本更新 57%

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

弹出式干扰揭示视频大语言模型中的事件袋行为

Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu, Dishant Zaveri, Khoa D. Doan, Kuan-Hao Huang

机构 * Texas A&M University(德克萨斯A&M大学) National Taiwan University(台湾国立大学) Stanford University(斯坦福大学) VinUniversity(文大学)

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

AI总结 通过插入无关广告片段,发现视频大语言模型常将不同片段的事件错误关联,表现出将视频视为事件集合而非时间序列的“事件袋”行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10932 2026-08-12 cs.CV cs.AI 新提交 57%

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

基于几何知识蒸馏的时序锚定组合式相机运动理解

Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo

机构 * The Hong Kong University of Science and Technology(香港科技大学) Tencent(腾讯)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 本研究提出CamDistill方法,通过几何知识蒸馏实现时序锚定的组合式相机运动理解,推出含4229个片段的CamChoreo基准,提升了相机运动识别的细粒度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08839 2026-08-11 cs.RO cs.CV 新提交 57%

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

SG-WAM:面向世界-动作模型的文本 grounded 与空间感知语义引导

Junjie He, Junfeng Li, Zhide Zhong, Haodong Yan, Ruixin Li, Yangyang Zheng, Jiaguan Zhu, Tianran Zhang, Yuqiao Du, Wen Chen, Shunbo Zhou, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Ola Dimensions(奥拉维度公司)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 针对现有世界-动作模型(WAM)存在的视频与指令语义不对齐问题,提出基于视觉-语言模型(VLM)的SG-WAM语义引导方法,通过注入文本接地且空间感知的语义前瞻,提升了WAM的指令遵循与操纵精度,经仿真和真实实验验证有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25465 2026-08-11 cs.CV cs.AI 版本更新 57%

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

EchoStyle: 通过反向数据合成解锁高保真视频风格化

Huaqiu Li, Jiahao Wang, Sijia Cai, Hualian Sheng, Bing Deng, Jieping Ye, Wenhan Luo

机构 * The Hong Kong University of Science and Technology(香港科技大学) Alibaba Group(阿里巴巴集团) Xi’an Jiaotong University(西安交通大学)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

AI总结 提出EchoStyle框架,通过构建视频到视频架构、自动反向合成数据集V-Style20k和初始跟随模式机制,解决视频风格化中的内容泄露、数据稀缺和长视频适应性问题,实现高质量任意长度视频风格化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07136 2026-08-10 cs.CV 版本更新 57%

MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding

MotionStrata:用于紧凑视频自编码的分层运动隐变量

Wenzhang Sun, Huaize Liu, Chunfeng Wang, Biao Gong, Hao Li, Changqing Zou

机构 * University of Chinese Academy of Sciences(中国科学院大学) Li Auto Zhejiang Lab(浙江实验室) State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 MotionStrata将运动预算分为全局与详细运动,在不增维度下实现分层表示,压缩时重建质量优于其他方法,为紧凑视频自编码提供新设计原则。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02330 2026-08-03 cs.CV cs.AI cs.LG 版本更新 57%

ActionParty: Multi-Subject Action Binding in Generative Video Games

ActionParty:生成视频游戏中的多主体动作绑定

Alexander Pondaven, Ziyi Wu, Igor Gilitschenski, Philip Torr, Sergey Tulyakov, Fabio Pizzati, Aliaksandr Siarohin

机构 * Snap Research(Snap研究院) University of Oxford(牛津大学) University of Toronto(多伦多大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 动作与事件理解 :video diffusion(abstract);分类 cs.CV

AI总结 本文提出ActionParty,一种可控制多主体的生成视频游戏世界模型,通过引入主体状态标记和空间偏置机制,提升动作关联准确性与身份一致性。

Comments ECCV 2026 - Project page: https://action-party.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28243 2026-07-31 cs.CV cs.AI 新提交 57%

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

EgoGenesis:基于在线锚定投影记忆与Action-3D RoPE的自我中心世界-动作建模

Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He, Kaibo Yin, Xiaobing Tu, Yinggui Wang, Jinkui Ren, Xiantao Zhang, Shijian Wang, Jinghong Liu, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团) Tianji KernalMind Co., Ltd.(天机芯智有限公司) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Renmin University of China(中国人民大学) The University of Tokyo(东京大学)

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

AI总结 本文提出EgoGenesis模拟器,通过在线锚定投影记忆与Action-3D RoPE合成自我中心操作视频,扩充训练数据,显著提升了真实机器人单臂、双臂任务的分布外操作成功率。

Comments project page: https://egogenesis.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13421 2026-07-16 cs.CV cs.AI 新提交 57%

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

ScanFocus:一种用于时空视频定位的从粗到细框架

Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang

机构 * Southeast University(东南大学)

专题命中 动作与事件理解 :long video(abstract);分类 cs.CV

AI总结 研究时空视频定位问题,提出ScanFocus从粗到细框架,利用统一融合编码器和轻量级模块生成粗略提议,再用语义引导时间聚合器恢复细节,实验表明该方法性能优于以往方法。

Comments this paper has already been accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏