arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6781 信号源:cs.CV, eess.IV, cs.MM

1. 视频扩散模型 1305 篇

2501.19252 2025-10-08 cs.CV 85%

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 视频扩散模型 :text-to-video(title,abstract);video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments Accepted to NeurIPS2025. Website: https://sites.google.com/view/t2v-dlbs and Code: https://github.com/shim0114/T2V-Diffusion-Search

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05096 2025-09-29 cs.CV 85%

Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers

Haosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu, Aiyue Chen, Jing Lin, Yiwu Yao, Chen Chen, Jingwen Leng, Yu Feng, Minyi Guo

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Qizhi Institute(上海启智研究院) Huawei Technologies Co.,Ltd(华为技术有限公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05323 2025-09-10 cs.AI cs.MM 85%

Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts

Adam Cole, Mick Grierson

机构 * University of the Arts London(伦敦艺术大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.MM

Comments 3rd international workshop on eXplainable AI for the Arts (XAIxArts) at the ACM Creativity and Cognition Conference June 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15894 2025-08-08 cs.CV 85%

RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers

Min Zhao, Guande He, Yixiao Chen, Hongzhou Zhu, Chongxuan Li, Jun Zhu

机构 * Dept. of Comp. Sci. \& Tech., BNRist Center, THU-Bosch ML Center, Tsinghua University. The University of Texas at Austin. Gaoling School of Artificial Intelligence Renmin University of China Beijing, China. Beijing Key Laboratory of Research on Large Models Engineering Research Center of Next-Generation Intelligent Search Pazhou Laboratory (Huangpu)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments ICML 2025. Project page: https://riflex-video.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06764 2025-07-25 cs.LG cs.CV 85%

History-Guided Video Diffusion

Kiwhan Song, Boyuan Chen, Max Simchowitz, Yilun Du, Russ Tedrake, Vincent Sitzmann

机构 * MIT(麻省理工学院) Harvard University(哈佛大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments ICML 2025. Project website: https://boyuan.space/history-guidance

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12847 2025-06-17 cs.GR cs.CV 85%

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Wei He, Jingdong Wang

机构 * Department of Computer Vision Technology(VIS), Baidu Inc.(计算机视觉技术系(VIS),百度公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments Technical report, 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03065 2025-06-04 cs.CV cs.AI cs.LG 85%

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers

Pengtao Chen, Xianfang Zeng, Maosen Zhao, Peng Ye, Mingzhu Shen, Wei Cheng, Gang Yu, Tao Chen

机构 * Fudan University(复旦大学) StepFun The Chinese University of Hong Kong(香港中文大学) Imperial College London(伦敦帝国理工学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08151 2025-05-20 cs.CV cs.LG 85%

Progressive Autoregressive Video Diffusion Models

Desai Xie, Zhan Xu, Yicong Hong, Hao Tan, Difan Liu, Feng Liu, Arie Kaufman, Yang Zhou

机构 * Stony Brook University(石溪大学) Adobe Research(Adobe研究)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments 15 pages, 7 figures. Code and video results are available at https://desaixie.github.io/pa-vdm/. v2: Accepted to CVPRW 2025. Updated figures, tables, notations, and text in all sections. Added comparison with more baseline methods, FVD metric results, user study, and discussion on parallel works

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18673 2025-05-07 cs.CV 85%

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin, Willi Menapace, Andrea Tagliasacchi, David B. Lindell, Sergey Tulyakov

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Snap Inc.(Snap公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments CVPR 2025; Project Page: https://snap-research.github.io/ac3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10783 2025-03-25 cs.CV 85%

Video Diffusion Transformers are In-Context Learners

Zhengcong Fei, Di Qiu, Debang Li, Changqian Yu, Mingyuan Fan

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12781 2025-03-25 cs.CV 85%

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace, Guocheng Qian, Michael Vasilkovsky, Hsin-Ying Lee, Chaoyang Wang, Jiaxu Zou, Andrea Tagliasacchi, David B. Lindell, Sergey Tulyakov

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments ICLR 2025; Project Page: https://snap-research.github.io/vd3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16983 2025-03-24 cs.CV cs.AI 85%

Enabling Versatile Controls for Video Diffusion Models

Xu Zhang, Hao Zhou, Haoming Qin, Xiaobin Lu, Jiaxing Yan, Guanzhong Wang, Zeyu Chen, Yi Liu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Codes and Supplementary Material: http://github.com/PaddlePaddle/PaddleMIX/tree/develop/ppdiffusers/examples/ppvctrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04012 2025-01-09 cs.MM cs.LG 85%

FlexCache: Flexible Approximate Cache System for Video Diffusion

Desen Sun, Henry Tian, Tim Lu, Sihang Liu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02741 2025-01-07 cs.CV 85%

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

Yunlong Yuan, Yuanfan Guo, Chunwei Wang, Hang Xu, Li Zhang

专题命中 视频扩散模型 :long video(title,abstract);video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14167 2024-12-19 cs.CV cs.AI cs.LG 85%

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Runtao Liu, Haoyu Wu, Zheng Ziqiang, Chen Wei, Yingqing He, Renjie Pi, Qifeng Chen

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03085 2024-12-05 cs.CV 85%

Mimir: Improving Video Diffusion Models for Precise Text Understanding

Shuai Tan, Biao Gong, Yutong Feng, Kecheng Zheng, Dandan Zheng, Shuwei Shi, Yujun Shen, Jingdong Chen, Ming Yang

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12831 2024-11-21 cs.CV 85%

Towards motion from video diffusion models

Paul Janson, Tiberiu Popa, Eugene Belilovsky

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop :Foundation Models for 3D Humans

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03150 2024-11-19 cs.CV cs.LG 85%

Video Diffusion Models: A Survey

Andrew Melnik, Michal Ljubljanac, Cong Lu, Qi Yan, Weiming Ren, Helge Ritter

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments https://github.com/ndrwmlnk/Awesome-Video-Diffusion-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03160 2024-10-07 cs.CV cs.LG 85%

Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach

Yaofang Liu, Yumeng Ren, Xiaodong Cun, Aitor Artola, Yang Liu, Tieyong Zeng, Raymond H. Chan, Jean-michel Morel

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments Code at https://github.com/Yaofang-Liu/FVDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10647 2024-09-17 cs.CV cs.AI cs.LG 85%

A Survey on Video Diffusion Models

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07537 2024-07-26 cs.CV 85%

FreeInit: Bridging Initialization Gap in Video Diffusion Models

Tianxing Wu, Chenyang Si, Yuming Jiang, Ziqi Huang, Ziwei Liu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Project page: https://tianxingwu.github.io/pages/FreeInit/ Code: https://github.com/TianxingWu/FreeInit

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10981 2024-06-18 cs.CV 85%

ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models

Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang, Jun Xiao

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments Code will be available at https://github.com/Dawn-LX/Causal-VideoGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02813 2024-04-10 cs.CV cs.AI 85%

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

Fengyuan Shi, Jiaxi Gu, Hang Xu, Songcen Xu, Wei Zhang, Limin Wang

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Accepted by CVPR 2024. Project page: https://bivdiff.github.io; GitHub repository: https://github.com/MCG-NJU/BIVDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13729 2024-04-05 cs.CV 85%

Hybrid Video Diffusion Models with 2D Triplane and 3D Wavelet Representation

Kihong Kim, Haneol Lee, Jihye Park, Seyeon Kim, Kwanghee Lee, Seungryong Kim, Jaejun Yoo

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments Project page is available at https://hxngiee.github.io/HVDM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10474 2024-03-27 cs.CV cs.GR cs.LG 85%

Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models

Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, Yogesh Balaji

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments ICCV 2023. Project webpage: https://research.nvidia.com/labs/dir/pyoco

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15169 2024-01-31 cs.CV 85%

FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling

Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, Ziwei Liu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);long video(abstract);分类 cs.CV

Comments ICLR 2024, Project Page: http://haonanqiu.com/projects/FreeNoise.html, Code Repo: https://github.com/AILab-CVC/FreeNoise

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06662 2023-12-12 cs.CV cs.AI cs.LG 85%

Photorealistic Video Generation with Diffusion Models

Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, José Lezama

专题命中 视频扩散模型 :video generation(title,abstract);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

Comments Project website https://walt-video-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03793 2023-12-08 cs.CV 85%

AnimateZero: Video Diffusion Models are Zero-Shot Image Animators

Jiwen Yu, Xiaodong Cun, Chenyang Qi, Yong Zhang, Xintao Wang, Ying Shan, Jian Zhang

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments Project Page: https://vvictoryuki.github.io/animatezero.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15127 2023-11-28 cs.CV 85%

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, Robin Rombach

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15103 2023-09-28 cs.CV 85%

LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, Yuwei Guo, Tianxing Wu, Chenyang Si, Yuming Jiang, Cunjian Chen, Chen Change Loy, Bo Dai, Dahua Lin, Yu Qiao, Ziwei Liu

专题命中 视频扩散模型 :video generation(title,abstract);text-to-video(abstract);long video(abstract);分类 cs.CV

Comments Project webpage: https://vchitect.github.io/LaVie-project/

详情

展开后加载摘要…

URL PDF HTML 收藏