Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
Multi-SpatialMLLM: 多模态大语言模型的多帧空间理解
机构 * FAIR, Meta(FAIR,Meta) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
AI总结 提出一种框架,通过整合深度感知、视觉对应和动态感知等基本空间技能,使多模态大语言模型具备多帧空间理解能力,并构建MultiSPA数据集和基准测试,模型Multi-SpatialMLLM在多项任务上取得显著提升。
Comments CVPR 2026 Camera Ready. 27 pages. Project page: https://runsenxu.com/projects/Multi-SpatialMLLM