Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale
任何3D场景都值得1000个标记:面向大规模场景生成的3D grounded表示
机构 * Westlake University(西湖大学) ; Afari Intelligent Drive(阿法里智能驾驶) ; Zhejiang University(浙江大学)
专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV
AI总结 本文提出在隐式3D潜在空间中直接生成3D场景,解决传统2D方法在3D空间扩展中的不足,通过3DRAE和3DDiT实现高效且一致的3D生成。
Comments Under Review. Project Page: https://wswdx.github.io/3DRAE