机构
*
The State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学复杂与关键软件环境国家重点实验室)
;
H3I, Beihang University(北京航空航天大学H3I)
;
China Mobile Information Technology Center(中国移动信息技术中心)
Any4D: Open-Prompt 4D Generation from Natural Language and Images
Any4D: 从自然语言和图像生成开放提示的4D生成
Hao Li, Qiao Sun
专题命中
视频生成
:video generation(abstract);分类 cs.CV
AI总结
本文提出Primitive Embodied World Models,通过限制视频生成时间范围,实现语言与视觉表示的细粒度对齐,降低学习复杂度,提升数据效率,并减少推理延迟,支持复杂任务的组合泛化。
CommentsThe authors identified issues in the 4D generation pipeline and evaluation that affect result validity. To ensure scientific accuracy, we will revise the methodology and experiments thoroughly before resubmitting. This version should not be cited or relied upon