Video Generation Models are General-Purpose Vision Learners
视频生成模型是通用视觉学习者
Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency
GenVid2Robot:通过刚性几何一致性从视频生成到机器人操作
Haohui Huang, Xi Yuan, Panpan Liao, Tao Teng, Chenguang Yang, Jing Guo, Yi Guo
机构
*
School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
;
University of Liverpool(利物浦大学)
;
Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)
;
State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学海洋地球科学国家重点实验室,自动化与智能感知学院)