MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
机构 * Apple(苹果公司)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments ICCV 2025
视觉与机器人
三维重建、NeRF、Gaussian Splatting、点云和空间智能。
机构 * Apple(苹果公司)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments ICCV 2025
机构 * School of Computer Science(计算机科学学院) ; University College Dublin(都柏林大学) ; School of Computing(计算机科学学院) ; Dublin City University(都柏林城市大学) ; School of Electronic Engineering(电子工程学院) ; Trinity College Dublin(都柏林三一学院)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments The benchmark MulSeT is available at https://huggingface.co/datasets/WanyueZhang/MulSeT
机构 * Zhejiang University(浙江大学) ; vivo ; Ant Group(蚂蚁集团)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments 21 pages, 12 figures. Accepted to ICCV 2025
机构 * Beijing Digital Native Digital City Research Center(北京数字原生数字城市研究中心) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 空间理解 :3D vision(title,abstract);分类 cs.CV
Comments Accepted by ICME2025
机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Waytous ; University of Technology Sydney(悉尼大学) ; Microsoft Research(微软研究院) ; IAIR, Xi’an Jiaotong University(西安交通大学IAIR) ; TikTok ; AIR, Tsinghua University(清华大学AIR)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
机构 * École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
专题命中 空间理解 :point cloud(title,abstract);分类 cs.CV
Comments Accepted for publication through the upcoming CVPR Workshop on open scene understanding with foundation models (OPENSUN3D)
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments ICRA 2025
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments Accepted by CVPR 2025
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
专题命中 空间理解 :3D vision(title,abstract);分类 cs.CV
Comments Accepted to 3DV 2025. Update 12/09/24: Change the benchmark name to UniQA-3D, add link to code
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments Accepted by ACL 2024 Main
专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV
Comments NAACL 2024
专题命中 空间理解 :3D vision(title);point cloud(abstract);分类 cs.RO
Comments 24 pages
专题命中 空间理解 :spatial understanding(title,abstract)
Comments Wordplay Workshop @ ACL 2024
GeoAnchor:通过潜在分解进行协作推理以实现3D空间理解
机构 * Shanghai Jiao Tong University(上海交通大学) ; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd(星辰通用人工智能实验室,中国电信人工智能技术(北京)有限公司) ; University of Science and Technology Beijing(北京科技大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
AI总结 针对从2D图像理解3D空间关系的挑战,提出GeoAnchor框架,通过分解3D空间信息为互补组件并结合协作训练策略,实现动态可解释推理,在复杂3D推理任务中优于现有技术。
Comments Accepted by ACM MM 2026
迷失于体积:CT-SpatialVQA基准用于评估3D医学视觉-语言模型的语义-空间理解
机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) ; New York University Abu Dhabi(纽约大学阿布扎克分校)
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
AI总结 本文提出CT-SpatialVQA基准,评估3D医学视觉-语言模型对3DCT数据的语义-空间理解能力,发现现有模型在语义-空间推理任务中表现不佳,需更深入整合体积证据以确保临床可靠性。
4DThinker: 用4D图像进行动态空间理解的思考
机构 * Tsinghua University, SIGS(清华大学 SIGS) ; Meituan(美团) ; The Chinese University of Hong Kong(香港中文大学) ; National University of Singapore(新加坡国立大学) ; LMMs-Lab(LMMs实验室) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
AI总结 提出4DThinker框架,通过动态潜在心理图像(在连续隐藏空间中模拟场景演化)增强视觉语言模型的动态空间推理能力,并引入无标注数据生成、动态图像微调及4D强化学习,在多个基准上超越强基线。
Comments 21 pages, 16 figures
OmniVLA-RL:具备空间理解与在线强化学习的视觉-语言-动作模型
机构 * AI Lab, Country Garden Services(国家花园服务人工智能实验室) ; Omni AI(奥米尼人工智能) ; VBot ; East China Normal University(东华大学)
专题命中 空间理解 :spatial understanding(title);分类 cs.RO
AI总结 本文提出OmniVLA-RL模型,通过混合Transformer架构整合推理、空间和动作专家,结合Flow-GSPO提升动作精度与训练稳定性,在LIBERO和LIBERO-Plus基准上表现优异,克服了现有VLA模型的局限。
无视位置,语言偏见:探测视觉语言编码器中中间层表征偏见以实现零样本语言基础空间理解
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) ; Samsung Electronics(三星电子)
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
AI总结 研究探讨了视觉语言编码器中间层中位置和语言相关信息的偏见,通过层间分析发现常规最终层多模态嵌入优先考虑全局语义对齐,导致视觉嵌入对位置线索敏感性低,多语言文本嵌入在共享空间中形成语言依赖的几何偏移,提出构建空间地图的方法提升零样本图像分割性能。
Comments 61 pages, 28 Figures, 15 Tables
VoxRep:通过体素表示增强2D视觉-语言模型的3D空间理解
机构 * Menlo Research(Menlo研究)
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
AI总结 本文提出VoxRep方法,通过将体素空间切分为2D切片并输入预训练的视觉-语言模型,实现对3D环境的高效语义理解。
Journal ref Proc. APSIPA ASC 2025, pp. 1464-1469
专题命中 空间理解 :spatial understanding(title);分类 cs.CV
Comments 14 pages, 11 figures
专题命中 空间理解 :3D vision(title);分类 cs.CV
Comments 10 pages, Keywords: design space exploration, machine learning, computer vision, SLAM, embedded systems, GPU, crowd-sourcing
Journal ref 31st IEEE International Parallel and Distributed Processing Symposium May 29 - June 2, 2017 Orlando, Florida USA
面向几何与纹理一致性的单图像引导模型生成与布局优化的3D场景生成
机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) ; Peng Cheng Laboratory(鹏城实验室) ; Harbin Institute of Technology(哈尔滨工业大学) ; Harbin Institute of Technology, Suzhou Research Institute(哈尔滨工业大学苏州研究院)
专题命中 空间理解 :point cloud(abstract);3D generation(abstract);分类 cs.CV、cs.GR
AI总结 本文提出一种三阶段框架,通过单图像引导的模型生成和空间布局优化,实现具有高几何准确性和纹理保真的3D场景生成。
Comments 14 pages, 9 figures, Project page: https://xdlbw.github.io/sing3d/
世界并非单一声场:为大型音频-语言模型启用空间理解
机构 * School of Intelligence Science and Technology(智能科学与技术学院)
专题命中 空间理解 :spatial understanding(title)
AI总结 本文提出TWNM框架,通过物理基础的FOA模拟和元数据引导训练,实现音频场景分析的三层次能力,提升空间音频语言理解的准确性和可审计性。
Comments 25 pages, 4 figures
网格空间理解:一个用于文本空间推理、具身场景和坐标结构的数据集
机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 空间理解 :spatial understanding(title)
AI总结 本文提出GSU数据集,用于评估LLM在导航、物体定位和结构组合任务中的空间推理能力,发现模型在具身代理参考框架和坐标列表识别上存在困难,且微调小模型可能达到前沿模型性能。
Comments preprint
SpatialText: 一种纯文本的认知基准,用于大型语言模型中的空间理解
机构 * Zhejiang University(浙江大学) ; School of Software Technology, Zhejiang University(软件技术学院,浙江大学)
专题命中 空间理解 :spatial understanding(title)
AI总结 SpatialText通过双源方法为大型语言模型提供纯文本空间认知基准,揭示其在空间推理中的系统性缺陷。