MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
MMPhysVideo: 通过联合多模态建模提升视频生成的物理合理性
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; StepFun(阶跃星辰) ; Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(多模态信息超级智能安全北京市重点实验室) ; School of Information Science and Technology, Shanghai Tech University(上海科技大学信息科学与技术学院)
专题命中 VLM训练与架构 :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 MMPhysVideo通过联合多模态建模提升视频生成的物理合理性,将感知线索统一为伪RGB格式,提出双向控制教师架构以减少跨模态干扰,并通过MMPhysPipe构建物理丰富的多模态数据集,提升物理合理性和视觉质量。
Comments Project Page: https://shubolin028.github.io/MMPhysVideo-Page