MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
MMPhysVideo: 通过联合多模态建模提升视频生成的物理合理性
Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
StepFun(阶跃星辰)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(多模态信息超级智能安全北京市重点实验室)
;
School of Information Science and Technology, Shanghai Tech University(上海科技大学信息科学与技术学院)
SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton
SymphonyGen:具有可控和声骨架的3D分层交响乐生成
Xuzheng He, Nan Nan, Zhilin Wang, Ziyue Kang, Zhuoru Mo, Ao Li, Yu Pan, Xiaobing Li, Feng Yu, Xiaohong Guan
机构
*
Department of AI Music and Music Information Technology, Central Conservatory of Music(中央音乐学院人工智能音乐与音乐信息科技系)
;
Frontier Institute of Science and Technology, and Interdisciplinary Research Center of Frontier Science and Technology, Xi’an Jiaotong University(西安交通大学前沿科学与技术研究院及交叉科学与技术 interdisciplinary research center)
;
University of Science and Technology of China(中国科学技术大学)
;
Shenzhen University(深圳大学)
DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
DriveCode: 针对基于LLM的自动驾驶的领域特定数值编码
Zhiye Wang, Yanbo Jiang, Rui Zhou, Bo Zhang, Fang Zhang, Zhenhua Xu, Yaqin Zhang, Jianqiang Wang
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
The School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)
;
The Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
DiDi, Beijing, China(滴滴出行)
;
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与移动性国家重点实验室)