SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
SpatiaLab: 视觉-语言模型能否在真实环境中进行空间推理?
机构 * Computational Intelligence and Operations Laboratory(计算智能与运筹实验室) ; Shahjalal University of Science and Technology(沙赫jalal科技大学) ; BRAC University(BRAC大学) ; North South University(北南大学) ; Monash University(墨尔本大学) ; Qatar Computing Research Institute(卡塔尔计算研究院)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.LG
AI总结 SpatiaLab通过真实场景下的空间推理任务评估视觉-语言模型的能力,揭示其在复杂空间关系、深度感知和3D几何方面的不足。
Comments Accepted to ICLR 2026 (https://openreview.net/forum?id=fWWUPOb0CT). 92 Pages. 42 Figures and 29 Tables
Journal ref ICLR 2026