BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension
BioVid: 具有生物行为语义理解的自回归视频生成
Tsung-Wei Pan, Jung-Hua Wang
机构
*
Department of Electrical Engineering, National Taiwan Ocean University(国立台湾海洋大学电子工程系)
;
AI research center, National Taiwan Ocean University(国立台湾海洋大学人工智能研究中心)
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shanghai Jiaotong University(上海交通大学)
;
Nanjing University(南京大学)
;
Beihang University(北京航空航天大学)
机构
*
Lica World(Lica世界)
;
San Francisco, United States of America(美国旧金山)
;
ICML’26 Workshop on Human-AI Co-Creativity, Seoul, South Korea(ICML’26 人类-人工智能协同创作研讨会,韩国首尔)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
VidCRAFT3: 面向图像到视频生成的相机、物体与光照控制
Sixiao Zheng, Zimian Peng, Yanpeng Zhou, Yi Zhu, Hang Xu, Xiangru Huang, Yanwei Fu
机构
*
School of Data Science, Fudan University(复旦大学数据科学学院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Zhejiang University(浙江大学)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
Westlake University(西湖大学)
;
School of Data Science and MOE Frontiers Center for Brain Science, Fudan University(复旦大学数据科学学院和脑科学前沿中心)
;
Fudan ISTBI–ZJNU Algorithm Centre for Brain-inspired Intelligence, Zhejiang Normal University(复旦大学-浙江师范大学脑启发智能算法中心)
机构
*
The University of British Columbia(不列颠哥伦比亚大学)
;
ETH Zürich(苏黎世联邦理工学院)
;
McMaster University(麦马斯特大学)
;
Vector Institute(向量研究所)
;
Canada CIFAR AI Chair(加拿大 CIFAR 人工智能主席)
机构
*
Zhipu AI, Beijing, China(智谱AI,北京,中国)
;
Tsinghua University, Beijing, China(清华大学,北京,中国)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai Qizhi Institute(上海启智研究院)
;
Huawei Technologies Co.,Ltd(华为技术有限公司)
;
Institute of Computing Technology, Chinese Academy of Science(中国科学院计算技术研究所)
Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva, Vladimir Arkhipkin, Vladimir Korviakov, Vladimir Polovnikov, Viacheslav Vasilev, Evelina Sidorova, Denis Dimitrov
CommentsReplacement note: This manuscript has been transferred from the CVPR format to the ECCV 2026 format, with the corresponding title and template updated accordingly. The technical content remains largely unchanged from the previous version. (Current version: 21 pages, 6 figures, and 3 tables.) Accepted by ECCV 2026