arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 2143 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 2143 篇

2405.18406 2025-10-06 cs.CV cs.AI cs.CL 57%

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 视频生成 :video diffusion(abstract);分类 cs.CV

Comments EMNLP 2025 main; The first two authors contribute equally. Project Page: https://raccoon-mllm-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10540 2025-10-03 cs.MA cs.CV 57%

AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation

Haoyuan Shi, Yunxin Li, Xinyu Chen, Longyue Wang, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Alibaba International Digital Commerce, Hangzhou(阿里巴巴国际数字商业)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02259 2025-10-03 cs.CV cs.AI 57%

VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

Mingzhe Zheng, Yongqi Xu, Haojian Huang, Xuran Ma, Yexin Liu, Wenjie Shu, Yatian Pang, Feilong Tang, Qifeng Chen, Harry Yang, Ser-Nam Lim

机构 * Hong Kong University of Science and Technology(香港科技大学) Peking University(北京大学) University of Hong Kong(香港大学) NUS(南洋理工大学) University of Central Florida(佛罗里达大学) Everlyn AI

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Code: https://github.com/DuNGEOnmassster/VideoGen-of-Thought.git; Webpage: https://cheliosoops.github.io/VGoT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10958 2025-10-02 cs.LG cs.AI cs.CV cs.NE cs.PF 57%

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen

机构 * Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,THBI实验室,清华-博世联合机器学习中心,清华大学) Institute for Interdisciplinary Information Sciences, Tsinghua University(交叉信息学院,清华大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments @inproceedings{zhang2024sageattention2, title={Sageattention2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization}, author={Zhang, Jintao and Huang, Haofeng and Zhang, Pengle and Wei, Jia and Zhu, Jun and Chen, Jianfei}, booktitle={International Conference on Machine Learning (ICML)}, year={2025} }

Journal ref Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 2025 (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14685 2025-10-02 cs.CV 57%

DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images

Kazuma Nagata, Naoshi Kaneko

机构 * Tokyo Denki University(东京电讯大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. v2: Added results on the subset used by the baseline for consistency; full test set results are also reported (Tables 1 and 2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24652 2025-09-30 cs.CV 57%

Learning Object-Centric Representations Based on Slots in Real World Scenarios

Adil Kaan Akan

机构 * Computer Science and Engineering(计算机科学与工程)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments PhD Thesis, overlap with arXiv:2507.20855 and arXiv:2501.15878

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23169 2025-09-30 cs.CV 57%

Sparse2Dense: A Keypoint-driven Generative Framework for Human Video Compression and Vertex Prediction

Bolin Chen, Ru-Ling Liao, Yan Ye, Jie Chen, Shanzhi Yin, Xinrui Ju, Shiqi Wang, Yibo Fan

机构 * Fudan University(复旦大学) DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华潘实验室) City University of Hong Kong(香港城市大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22432 2025-09-29 cs.CV 57%

Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy

Yuhao Liu, Tengfei Wang, Fang Liu, Zhenwei Wang, Rynson W. H. Lau

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 视频生成 :video diffusion(abstract);分类 cs.CV

Comments Accepted by Siggraph Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02045 2025-09-26 cs.GR cs.CV 57%

Generating 360° Video is What You Need For a 3D Scene

Zhaoyang Zhang, Yannick Hold-Geoffroy, Miloš Hašan, Ziwen Chen, Fujun Luan, Julie Dorsey, Yiwei Hu

机构 * Yale University(耶鲁大学) Adobe Research(Adobe研究) Oregon State University(俄勒冈州立大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments SIGGRAPH Asia 2025. Project Page: https://zhaoyangzh.github.io/projects/worldprompter/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09595 2025-09-18 cs.CV 57%

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://klingavatar.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09632 2025-09-09 cs.CV cs.AI 57%

Preacher: Paper-to-Video Agentic System

Jingwei Liu, Ling Yang, Hao Luo, Fan Wang, Hongyan Li, Mengdi Wang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) Hupan Lab(虎扑实验室) National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01028 2025-09-04 cs.CV 57%

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

Zixin Zhu, Kevin Duarte, Mamshad Nayeem Rizve, Chengyuan Xu, Ratheesh Kalarot, Junsong Yuan

机构 * University at Buffalo(布法罗大学) Adobe Inc. (ASML)(Adobe公司)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12278 2025-09-04 cs.CV 57%

Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Content

Rohit Kundu, Hao Xiong, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury

机构 * Google, Mountain View, USA(谷歌(Mountain View, USA)) University of California, Riverside(加州大学河滨分校)

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01232 2025-09-03 cs.CV 57%

FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework

Lingzhou Mu, Qiang Wang, Fan Jiang, Mengchao Wang, Yaqi Fan, Mu Xu, Kai Zhang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments https://fantasy-amap.github.io/fantasy-hsi/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00403 2025-09-03 cs.CV 57%

DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective

Yushuo Chen, Ruizhi Shao, Youxin Pang, Hongwen Zhang, Xinyi Wu, Rihui Wu, Yebin Liu

机构 * Tsinghua University(清华大学) Beijing Normal University(北京师范大学) Honor Device Co., Ltd(荣誉设备有限公司)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22141 2025-08-28 cs.CV cs.AI 57%

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing

Guanwen Feng, Zhiyuan Ma, Yunan Li, Jiahao Yang, Junwei Jing, Qiguang Miao

机构 * Xi’an Key Laboratory of Big Data and Intelligent Vision, Xidian University, Xi’an 710071, China(西安大数据与智能视觉重点实验室,西安电子科技大学,西安710071,中国) Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University, Xi’an 710071, China(协同智能系统重点实验室,教育部,西安电子科技大学,西安710071,中国) School of Computer Science and Technology, Xidian University, Xi’an 710071, China(计算机科学与技术学院,西安电子科技大学,西安710071,中国)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15774 2025-08-22 cs.CV 57%

CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

Haonan Qiu, Ning Yu, Ziqi Huang, Paul Debevec, Ziwei Liu

机构 * Nanyang Technological University(南洋理工大学) Netflix Eyeline Studios

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments CineScale is an extended work of FreeScale (ICCV 2025). Project Page: https://eyeline-labs.github.io/CineScale/, Code Repo: https://github.com/Eyeline-Labs/CineScale

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15773 2025-08-22 cs.CV cs.GR cs.LG 57%

Scaling Group Inference for Diverse and High-Quality Generation

Gaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang, Kfir Aberman, Srinivasa Narasimhan, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) Snap Research(Snap研究公司) Tel Aviv University(特拉维夫大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project website: https://www.cs.cmu.edu/~group-inference, GitHub: https://github.com/GaParmar/group-inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12415 2025-08-22 cs.CV 57%

TiP4GEN: Text to Immersive Panorama 4D Scene Generation

Ke Xing, Hanwen Liang, Dejia Xu, Yuyang Yin, Konstantinos N. Plataniotis, Yao Zhao, Yunchao Wei

机构 * Institute of Information Science, Beijing Jiaotong University(信息科学研究院,北京交通大学) Visual Intelligence + X International Joint Laboratory(视觉智能+X国际联合实验室) University of Toronto(多伦多大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted In Proceedings of the 33rd ACM International Conference on Multimedia (MM' 25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14465 2025-08-21 cs.CV 57%

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

Weitao Wang, Zichen Wang, Hongdeng Shen, Yulei Lu, Xirui Fan, Suhui Wu, Jun Zhang, Haoqian Wang, Hao Zhang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13797 2025-08-20 cs.GR cs.CV 57%

Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing

Feng-Lin Liu, Shi-Yang Li, Yan-Pei Cao, Hongbo Fu, Lin Gao

机构 * Institute of Computing Technology, Chinese Academy of Sciences, China(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Hong Kong University of Science(香港科学大学)

专题命中 视频生成 :video diffusion(abstract);分类 cs.CV

Comments SIGGRAPH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13013 2025-08-19 cs.CV 57%

EgoTwin: Dreaming Body and View in First Person

Jingqiao Xiu, Fangzhou Hong, Yicong Li, Mengze Li, Wentao Wang, Sirui Han, Liang Pan, Ziwei Liu

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11049 2025-08-18 cs.RO cs.CV 57%

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

Kelin Yu, Sheng Zhang, Harshit Soora, Furong Huang, Heng Huang, Pratap Tokekar, Ruohan Gao

机构 * University of Maryland, College Park(马里兰大学 College Park分校)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12341 2025-08-14 cs.MM 57%

Multimodal LLM-based Query Paraphrasing for Video Search

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Sheng-Hua Zhong, Xiong-Yong Wei, Qing Li

专题命中 视频生成 :text-to-video(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06080 2025-08-11 cs.CV 57%

DreamVE: Unified Instruction-based Image and Video Editing

Bin Xia, Jiyang Liu, Yuechen Zhang, Bohao Peng, Ruihang Chu, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) ByteDance Inc(字节跳动公司) HKUST(香港理工大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06364 2025-08-11 cs.CV cs.GR 57%

Generative Video Bi-flow

Chen Liu, Tobias Ritschel

机构 * University College London(伦敦大学学院)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments ICCV 2025. Project Page at https://ryushinn.github.io/ode-video

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00701 2025-08-06 cs.CV cs.AI 57%

D3: Training-Free AI-Generated Video Detection Using Second-Order Features

Chende Zheng, Ruiqi suo, Chenhao Lin, Zhengyu Zhao, Le Yang, Shuai Liu, Minghui Yang, Cong Wang, Chao Shen

机构 * Xi’an Jiaotong University(西安交通大学) Guangdong OPPO Mobile Communications Co., Ltd.(广东OPPO移动通信有限公司) City University of Hong Kong(香港城市大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00299 2025-08-04 cs.CV cs.AI cs.RO 57%

Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence

Danzhen Fu, Jiagao Hu, Daiguo Zhou, Fei Wang, Zepeng Wang, Wenhua Liao

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments ICCV 2025 Workshop (HiGen)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12024 2025-07-31 cs.CV 57%

SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering

Byeongjun Park, Hyojun Go, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, Changick Kim

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project page: https://byeongjun-park.github.io/SteerX/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09631 2025-07-30 cs.GR eess.IV 57%

V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video

Jianqi Chen, Biao Zhang, Xiangjun Tang, Peter Wonka

专题命中 视频生成 :video generation(abstract);分类 eess.IV

Comments Accepted by ICCV 2025. Project page: https://windvchen.github.io/V2M4/

详情

展开后加载摘要…

URL PDF HTML 收藏