MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
MMPhysVideo: 通过联合多模态建模提升视频生成的物理合理性
Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
StepFun(阶跃星辰)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(多模态信息超级智能安全北京市重点实验室)
;
School of Information Science and Technology, Shanghai Tech University(上海科技大学信息科学与技术学院)
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
Prompting-MammAlps:用于相机陷阱数据的细粒度文本到视频检索
Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis
机构
*
Ecole Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)
;
Max Planck Institute of Animal Behavior(马克斯·普朗克动物行为研究所)
;
University of Konstanz(康斯坦茨大学)
Video Generation Models are General-Purpose Vision Learners
视频生成模型是通用视觉学习者
Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora
CourseBlueprint:基于课程语料库的自适应教学视频生成的结构化流程
Md Zabirul Islam, Md Motaleb Hossen Manik, Ge Wang
机构
*
Department of Computer Science, Rensselaer Polytechnic Institute(计算机科学系,拉特格斯理工学院)
;
Department of Biomedical Engineering, Rensselaer Polytechnic Institute(生物医学工程系,拉特格斯理工学院)