arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 497 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 497 篇

2503.13111 2025-09-09 cs.CV cs.CL cs.LG 79%

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch

机构 * Apple(苹果公司)

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02175 2025-09-05 cs.CV cs.AI cs.CL cs.LG 79%

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

Nils Hoehing, Mayug Maniparambil, Ellen Rushe, Noel E. O'Connor, Anthony Ventresque

机构 * School of Computer Science(计算机科学学院) University College Dublin(都柏林大学) School of Computing(计算机科学学院) Dublin City University(都柏林城市大学) School of Electronic Engineering(电子工程学院) Trinity College Dublin(都柏林三一学院)

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02359 2025-09-03 cs.CV 79%

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Wanyue Zhang, Yibin Huang, Yangbin Xu, JingJing Huang, Helu Zhi, Shuo Ren, Wang Xu, Jiajun Zhang

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments The benchmark MulSeT is available at https://huggingface.co/datasets/WanyueZhang/MulSeT

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13195 2025-08-26 cs.CV 79%

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

Gaoyang Zhang, Bingtao Fu, Qingnan Fan, Qi Zhang, Runxing Liu, Hong Gu, Huaqi Zhang, Xinguo Liu

机构 * Zhejiang University(浙江大学) vivo Ant Group(蚂蚁集团)

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments 21 pages, 12 figures. Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16524 2025-07-23 cs.CV cs.AI 79%

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models

Xiaoyan Wang, Zeju Li, Yifan Xu, Jiaxing Qi, Zhifei Yang, Ruifei Ma, Xiangde Liu, Chao Zhang

机构 * Beijing Digital Native Digital City Research Center(北京数字原生数字城市研究中心) The Chinese University of Hong Kong(香港中文大学)

专题命中 空间理解 :3D vision(title,abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13112 2025-05-28 cs.CV 79%

SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Xianda Guo, Ruijun Zhang, Yiqun Duan, Yuhang He, Dujun Nie, Wenke Huang, Chenming Zhang, Shuai Liu, Hao Zhao, Long Chen

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Waytous University of Technology Sydney(悉尼大学) Microsoft Research(微软研究院) IAIR, Xi’an Jiaotong University(西安交通大学IAIR) TikTok AIR, Tsinghua University(清华大学AIR)

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13590 2025-04-21 cs.CV cs.AI 79%

HAECcity: Open-Vocabulary Scene Understanding of City-Scale Point Clouds with Superpoint Graph Clustering

Alexander Rusnak, Frédéric Kaplan

机构 * École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 空间理解 :point cloud(title,abstract);分类 cs.CV

Comments Accepted for publication through the upcoming CVPR Workshop on open scene understanding with foundation models (OPENSUN3D)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03164 2025-04-08 cs.CV cs.AI 79%

NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Kexin Tian, Jingrui Mao, Yunlong Zhang, Jiwan Jiang, Yang Zhou, Zhengzhong Tu

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06613 2025-04-07 cs.CV 79%

3D Spatial Understanding in MLLMs: Disambiguation and Evaluation

Chun-Peng Chang, Alain Pagani, Didier Stricker

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00379 2025-04-02 cs.CV 79%

MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving

Zhiyuan Zhang, Xiaofan Li, Zhihao Xu, Wenjie Peng, Zijian Zhou, Miaojing Shi, Shuangping Huang

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13642 2025-03-20 cs.CV 79%

SpatialBot: Precise Spatial Understanding with Vision Language Models

Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, Bo Zhao

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03878 2025-03-04 cs.CV 79%

SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models

Yue Zhang, Zhiyang Xu, Ying Shen, Parisa Kordjamshidi, Lifu Huang

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10799 2024-12-10 cs.CV 79%

Towards Foundation Models for 3D Vision: How Close Are We?

Yiming Zuo, Karhan Kayan, Maggie Wang, Kevin Jeon, Jia Deng, Thomas L. Griffiths

专题命中 空间理解 :3D vision(title,abstract);分类 cs.CV

Comments Accepted to 3DV 2025. Update 12/09/24: Change the benchmark name to UniQA-3D, add link to code

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15308 2024-06-12 cs.CV cs.LG 79%

SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding

Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemulapalli, Mehrdad Farajtabar, Sachin Mehta, Mohammad Rastegari, Oncel Tuzel, Hadi Pouransari

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05756 2024-06-11 cs.AI cs.CL cs.CV cs.MM 79%

EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang, Zhongyu Wei

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments Accepted by ACL 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13594 2024-04-23 cs.CV cs.AI 79%

Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers

Georgios Pantazopoulos, Alessandro Suglia, Oliver Lemon, Arash Eshghi

专题命中 空间理解 :spatial understanding(title,abstract);分类 cs.CV

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.04970 2017-10-16 cs.RO 79%

Transfer of Tool Affordance and Manipulation Cues with 3D Vision Data

Paulo Abelha, Frank Guerin

专题命中 空间理解 :3D vision(title);point cloud(abstract);分类 cs.RO

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04648 2024-08-12 cs.CL cs.AI cs.IR 78%

PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models

Alexey Tikhonov

专题命中 空间理解 :spatial understanding(title,abstract)

Comments Wordplay Workshop @ ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13454 2026-07-29 cs.CV cs.AI 版本更新 74%

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

GeoAnchor:通过潜在分解进行协作推理以实现3D空间理解

Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, Zhiyu Lin, Ye Yuan, Zhongjiang He, Yu Yu, Hao Sun

机构 * Shanghai Jiao Tong University(上海交通大学) Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd(星辰通用人工智能实验室,中国电信人工智能技术(北京)有限公司) University of Science and Technology Beijing(北京科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

AI总结 针对从2D图像理解3D空间关系的挑战,提出GeoAnchor框架,通过分解3D空间信息为互补组件并结合协作训练策略,实现动态可解释推理,在复杂3D推理任务中优于现有技术。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08787 2026-06-23 cs.CV 版本更新 74%

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

迷失于体积:CT-SpatialVQA基准用于评估3D医学视觉-语言模型的语义-空间理解

Mashrafi Monon, Umaima Rahman, Asif Hanif, Numan Saeed, Mohammad Yaqub

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) New York University Abu Dhabi(纽约大学阿布扎克分校)

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

AI总结 本文提出CT-SpatialVQA基准,评估3D医学视觉-语言模型对3DCT数据的语义-空间理解能力,发现现有模型在语义-空间推理任务中表现不佳,需更深入整合体积证据以确保临床可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05997 2026-05-25 cs.CV 74%

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

4DThinker: 用4D图像进行动态空间理解的思考

Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xiang An, Bo Li, Xin Xie, ZiDong Wang, Mingze Sun, Shuang Chen, Hongyu Li, Xiaobin Hu, Ruqi Huang

机构 * Tsinghua University, SIGS(清华大学 SIGS) Meituan(美团) The Chinese University of Hong Kong(香港中文大学) National University of Singapore(新加坡国立大学) LMMs-Lab(LMMs实验室) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

AI总结 提出4DThinker框架,通过动态潜在心理图像(在连续隐藏空间中模拟场景演化)增强视觉语言模型的动态空间推理能力,并引入无标注数据生成、动态图像微调及4D强化学习,在多个基准上超越强基线。

Comments 21 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17706 2026-04-27 cs.RO 74%

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

OmniVLA-RL:具备空间理解与在线强化学习的视觉-语言-动作模型

Haoxiang Jie, Yaoyuan Yan, Xiangyu Wei, Kailin Wang, Hongjie Yan, Zhiyou Heng, Daocheng Chen

机构 * AI Lab, Country Garden Services(国家花园服务人工智能实验室) Omni AI(奥米尼人工智能) VBot East China Normal University(东华大学)

专题命中 空间理解 :spatial understanding(title);分类 cs.RO

AI总结 本文提出OmniVLA-RL模型,通过混合Transformer架构整合推理、空间和动作专家,结合Flow-GSPO提升动作精度与训练稳定性,在LIBERO和LIBERO-Plus基准上表现优异,克服了现有VLA模型的局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23098 2026-03-20 cs.CV cs.AI 74%

Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding

无视位置,语言偏见:探测视觉语言编码器中中间层表征偏见以实现零样本语言基础空间理解

Na Min An, Inha Kang, Minhyun Lee, Hyunjung Shim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) Samsung Electronics(三星电子)

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

AI总结 研究探讨了视觉语言编码器中间层中位置和语言相关信息的偏见,通过层间分析发现常规最终层多模态嵌入优先考虑全局语义对齐,导致视觉嵌入对位置线索敏感性低,多语言文本嵌入在共享空间中形成语言依赖的几何偏移,提出构建空间地图的方法提升零样本图像分割性能。

Comments 61 pages, 28 Figures, 15 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21214 2025-12-03 cs.CV cs.CL 74%

VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation

VoxRep:通过体素表示增强2D视觉-语言模型的3D空间理解

Alan Dao, Norapat Buppodom

机构 * Menlo Research(Menlo研究)

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

AI总结 本文提出VoxRep方法,通过将体素空间切分为2D切片并输入预训练的视觉-语言模型,实现对3D环境的高效语义理解。

Journal ref Proc. APSIPA ASC 2025, pp. 1464-1469

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18270 2024-12-04 cs.CV 74%

Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents

Joongwon Chae, Zhenyu Wang, Lian Zhang, Dongmei Yu, Peiwu Qin

专题命中 空间理解 :spatial understanding(title);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.00505 2017-03-23 cs.CV cs.DC cs.LG cs.PF 74%

Algorithmic Performance-Accuracy Trade-off in 3D Vision Applications Using HyperMapper

Luigi Nardi, Bruno Bodin, Sajad Saeedi, Emanuele Vespa, Andrew J. Davison, Paul H. J. Kelly

专题命中 空间理解 :3D vision(title);分类 cs.CV

Comments 10 pages, Keywords: design space exploration, machine learning, computer vision, SLAM, embedded systems, GPU, crowd-sourcing

Journal ref 31st IEEE International Parallel and Distributed Processing Symposium May 29 - June 2, 2017 Orlando, Florida USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14841 2026-02-18 cs.GR cs.CV 73%

Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization

面向几何与纹理一致性的单图像引导模型生成与布局优化的3D场景生成

Xiang Tang, Ruotong Li, Xiaopeng Fan

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Peng Cheng Laboratory(鹏城实验室) Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology, Suzhou Research Institute(哈尔滨工业大学苏州研究院)

专题命中 空间理解 :point cloud(abstract);3D generation(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种三阶段框架,通过单图像引导的模型生成和空间布局优化,实现具有高几何准确性和纹理保真的3D场景生成。

Comments 14 pages, 9 figures, Project page: https://xdlbw.github.io/sing3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02954 2026-05-12 cs.SD cs.AI 71%

The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models

世界并非单一声场:为大型音频-语言模型启用空间理解

Yuhuan You, Lai Wei, Xihong Wu, Tianshu Qu

机构 * School of Intelligence Science and Technology(智能科学与技术学院)

专题命中 空间理解 :spatial understanding(title)

AI总结 本文提出TWNM框架,通过物理基础的FOA模拟和元数据引导训练,实现音频场景分析的三层次能力,提升空间音频语言理解的准确性和可审计性。

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17333 2026-03-19 cs.CL 71%

Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures

网格空间理解:一个用于文本空间推理、具身场景和坐标结构的数据集

Risham Sidhu, Julia Hockenmaier

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 空间理解 :spatial understanding(title)

AI总结 本文提出GSU数据集,用于评估LLM在导航、物体定位和结构组合任务中的空间推理能力,发现模型在具身代理参考框架和坐标列表识别上存在困难,且微调小模型可能达到前沿模型性能。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03002 2026-03-04 cs.AI 71%

SpatialText: A Pure-Text Cognitive Benchmark for Spatial Understanding in Large Language Models

SpatialText: 一种纯文本的认知基准,用于大型语言模型中的空间理解

Peiyao Jiang, Zequn Qin, Xi Li

机构 * Zhejiang University(浙江大学) School of Software Technology, Zhejiang University(软件技术学院,浙江大学)

专题命中 空间理解 :spatial understanding(title)

AI总结 SpatialText通过双源方法为大型语言模型提供纯文本空间认知基准,揭示其在空间推理中的系统性缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏