Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
超人:统一骨骼与视觉用于人体运动感知与生成
机构 * State Key Laboratory of General Artificial Intelligence, Peking University, Shenzhen Graduate School(通用人工智能国家重点实验室,北京大学深圳研究生院) ; Sony R&D Center(索尼研发中心) ; Nanyang Technological University(南洋理工大学)
专题命中 多模态生成 :MLLM(summary_cn,abstract);cross-modal(abstract);分类 cs.CV
AI总结 针对人体运动分析范式碎片化问题,提出Superman统一框架,通过视觉引导运动分词器创建跨模态运动词汇,训练单一MLLM架构处理多任务,实验表明该方法在运动任务中性能领先或具竞争力,为运动分析提供高效可扩展路径。