Universal Skeleton Understanding via Differentiable Rendering and MLLMs
通过可微渲染和大语言模型实现通用骨架理解
机构 * State Key Laboratory of General Artificial Intelligence, Peking University, Shenzhen Graduate School, China(人工智能通用基础理论国家重点实验室,北京大学深圳研究生院,中国) ; Tencent(腾讯) ; Nanjing University of Aeronautics(南京航空航天大学)
专题命中 图文多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 本文提出SkeletonLLM,通过可微渲染将任意骨架序列转换为大语言模型的视觉模态,实现通用骨架理解,同时引入协同训练策略提升推理能力,展示了在开放词汇动作识别中的强泛化能力,并扩展到异构骨架格式的运动描述和问答任务。
Comments Accepted by ICML 2026