arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1298 信号源:cs.CV, eess.IV, cs.MM

1. 视频扩散模型 1298 篇

2506.08009 2025-11-11 cs.CV cs.AI cs.LG 83%

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, Eli Shechtman

机构 * Adobe Research(Adobe研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments NeurIPS 2025 spotlight. Project website: http://self-forcing.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10501 2025-10-31 cs.CV cs.LG 83%

OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models

Mathis Koroglu, Hugo Caselles-Dupré, Guillaume Jeanneret Sanmiguel, Matthieu Cord

机构 * Obvious Research ISIR - Sorbonne University(ISIR - 索邦大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments 8 pages, 1 supplementary page, 9 figures

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025, pp. 6225-6235

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19022 2025-10-23 cs.CV 83%

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

Aritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amirhossein Habibian, Mohsen Ghafoorian

机构 * University of Amsterdam(阿姆斯特丹大学) Qualcomm AI Research(高通AI研究)

专题命中 视频扩散模型 :video diffusion(title,abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04126 2025-10-20 cs.CV cs.AI 83%

Multi-identity Human Image Animation with Structural Video Diffusion

Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo, Dahua Lin, Tianfan Xue, Bo Dai

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments ICCV 2025 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14179 2025-10-17 cs.CV cs.AI 83%

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip, Ahmet Levent Taşel, Yiwei Zhao, Ryan Burgert, Mingming He, Oliver Hermann, Oliver Pilarski, Rahul Garg, Paul Debevec, Ning Yu

机构 * Eyeline Labs(Eyeline实验室)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07944 2025-10-17 cs.CV 83%

CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving

Tianrui Zhang, Yichen Liu, Zilin Guo, Yuxin Guo, Jingcheng Ni, Chenjing Ding, Dan Xu, Lewei Lu, Zehuan Wu

机构 * Sensetime Research(商汤科技研究院) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09789 2025-10-17 cs.CV cs.AI 83%

On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise

Chao Liu, Arash Vahdat

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12626 2025-10-16 cs.CV 83%

Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models

Lvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein, Maneesh Agrawala

机构 * Stanford University(斯坦福大学) MIT(麻省理工学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments https://github.com/lllyasviel/FramePack

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12785 2025-10-15 cs.CV cs.AI cs.GR 83%

MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars

Felix Taubner, Ruihang Zhang, Mathieu Tuli, Sherwin Bahmani, David B. Lindell

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Vector Institute Canada(加拿大向量研究所) University of Toronto Canada(多伦多大学加拿大)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10670 2025-10-14 cs.CV 83%

AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes

Yu Li, Menghan Xia, Gongye Liu, Jianhong Bai, Xintao Wang, Conglang Zhang, Yuxuan Lin, Ruihang Chu, Pengfei Wan, Yujiu Yang

机构 * Tsinghua University(清华大学) HUST(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) HKUST(香港科技大学) Zhejiang University(浙江大学) Wuhan University(武汉大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24695 2025-10-14 cs.CV cs.AI 83%

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, Haozhe Liu, Hongwei Yi, Hao Zhang, Muyang Li, Yukang Chen, Han Cai, Sanja Fidler, Ping Luo, Song Han, Enze Xie

机构 * NVIDIA HKU(香港大学) MIT(麻省理工学院) THU(清华大学) PKU(北京大学) KAUST(国王 Abdullah 基础研究科学研究院)

专题命中 视频扩散模型 :video generation(title,abstract);long video(abstract);分类 cs.CV

Comments 21 pages, 15 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18809 2025-10-14 cs.CV 83%

VORTA: Efficient Video Diffusion via Routing Sparse Attention

Wenhao Sun, Rong-Cheng Tu, Yifu Ding, Zhao Jin, Jingyi Liao, Shunyu Liu, Dacheng Tao

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. The code is available at https://github.com/wenhao728/VORTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03517 2025-10-13 cs.CV 83%

DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models

Ziyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Ashkan Mirzaei, Igor Gilitschenski, Sergey Tulyakov, Aliaksandr Siarohin

机构 * Snap Research University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 视频扩散模型 :video diffusion(title,abstract);text-to-video(abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight. Project page: https://snap-research.github.io/DenseDPO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07345 2025-10-10 q-bio.QM cs.AI eess.IV 83%

Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model

Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal, Omid Mohareri

机构 * a Department of Translational Surgical Oncology, NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden,HZDR, Germany(a 转化外科肿瘤学部,NCT/UCC 德累斯顿,DKFZ、医学院和卡尔·古斯塔夫·卡尔斯医院、德累斯顿技术大学、HZDR 的联合体,德国) b Intuitive Surgical, Inc., Sunnyvale, CA, United States(b 直觉手术公司,美国加利福尼亚州 Sunnyvale)

专题命中 视频扩散模型 :video diffusion(title,abstract);video understanding(abstract);分类 eess.IV

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07190 2025-10-09 cs.CV 83%

MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis

Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang, Zhengwentai Sun, Jiahao Chang, Xiaodong Cun, Wensen Feng, Xiaoguang Han

机构 * Great Bay University(大亚湾大学) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Future Networks of Intelligence(广东省未来网络智能化重点实验室)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted by SIGGRAPH Asia 2025 conference track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21653 2025-10-08 cs.CV 83%

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

Ke Zhang, Cihan Xiao, Jiacong Xu, Yiqun Mei, Vishal M. Patel

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26025 2025-10-01 cs.CV 83%

PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution

Shian Du, Menghan Xia, Chang Liu, Xintao Wang, Jing Wang, Pengfei Wan, Di Zhang, Xiangyang Ji

机构 * Tsinghua University(清华大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Beijing Institute of Technology(北京理工大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24997 2025-09-30 cs.CV 83%

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

Yuyang Yin, HaoXiang Guo, Fangfu Liu, Mengyu Wang, Hanwen Liang, Eric Li, Yikai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei

机构 * Beijing Jiaotong University(北京交通大学) Skywork AI Tsinghua University(清华大学) University of Toronto(多伦多大学) Beijing Normal University(北京师范大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project page: \url{https://yuyangyin.github.io/PanoWorld-X/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21839 2025-09-30 cs.CV cs.AI 83%

DiTraj: training-free trajectory control for video diffusion transformer

Cheng Lei, Jiayu Zhang, Yue Ma, Xinyu Wang, Long Chen, Liang Tang, Yiqiang Yan, Fei Su, Zhicheng Zhao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Lenovo(联想) HKUST(香港科技大学) Tsinghua University(清华大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07772 2025-09-25 cs.CV 83%

From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

Tianwei Yin, Qiang Zhang, Richard Zhang, William T. Freeman, Fredo Durand, Eli Shechtman, Xun Huang

机构 * MIT(麻省理工学院) Adobe(Adobe公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://causvid.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03485 2025-09-24 cs.CV 83%

LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation

Lianwei Yang, Haokun Lin, Tianchen Zhao, Yichen Wu, Hongyu Zhu, Ruiqi Xie, Zhenan Sun, Yu Wang, Qingyi Gu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 视频扩散模型 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09547 2025-09-12 cs.CV cs.AI 83%

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

Dohun Lee, Hyeonho Jeong, Jiwook Kim, Duygu Ceylan, Jong Chul Ye

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 17 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07484 2025-09-10 cs.CV 83%

LINR Bridge: Vector Graphic Animation via Neural Implicits and Video Diffusion Priors

Wenshuo Gao, Xicheng Lan, Luyao Zhang, Shuai Yang

机构 * Wangxuan Institute of Computer Technology, State Key Laboratory of Multimedia Information Processing, \ University, Beijing, China Wangxuan Institute of Computer Technology, State Key Laboratory of Multimedia Information Processing, Peking University, Beijing, China

专题命中 视频扩散模型 :video diffusion(title,abstract);text-to-video(abstract);分类 cs.CV

Comments 5 pages, ICIPW 2025, Website: https://gaowenshuo.github.io/LINR-bridge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03254 2025-08-06 cs.CV cs.AI 83%

V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models

Jisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park, Sooyeon Park, Youngjae Yu

机构 * Yonsei University(延世大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments ICCV2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01698 2025-08-05 cs.CV 83%

Versatile Transition Generation with Image-to-Video Diffusion

Zuhao Yang, Jiahui Zhang, Yingchen Yu, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22360 2025-07-31 cs.CV cs.AI 83%

GVD: Guiding Video Diffusion Model for Scalable Video Distillation

Kunyang Li, Jeffrey A Chan Santiago, Sarinda Dhanesh Samarasinghe, Gaowen Liu, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Cisco Research(思科研究)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07982 2025-07-11 cs.CV cs.AI 83%

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 18 pages, project page: https://GeometryForcing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02860 2025-07-04 cs.CV 83%

Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching

Xin Zhou, Dingkang Liang, Kaijin Chen, Tianrui Feng, Xiwu Chen, Hongkai Lin, Yikang Ding, Feiyang Tan, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MEGVII Technology(美科七科技) University of Hong Kong(香港大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments The code is made available at https://github.com/H-EmbodVis/EasyCache. Project page: https://h-embodvis.github.io/EasyCache/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23858 2025-07-01 cs.CV 83%

VMoBA: Mixture-of-Block Attention for Video Diffusion Models

Jianzong Wu, Liang Hou, Haotian Yang, Xin Tao, Ye Tian, Pengfei Wan, Di Zhang, Yunhai Tong

机构 * Peking University(北京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Code is at https://github.com/KwaiVGI/VMoBA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17220 2025-06-24 cs.CV 83%

Emergent Temporal Correspondences from Video Diffusion Transformers

Jisu Nam, Soowon Son, Dahyun Chung, Jiyoung Kim, Siyoon Jin, Junhwa Hur, Seungryong Kim

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project page is available at https://cvlab-kaist.github.io/DiffTrack

详情

展开后加载摘要…

URL PDF HTML 收藏