arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 497 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 497 篇

2310.14540 2024-04-16 cs.CL cs.AI 71%

Evaluating Spatial Understanding of Large Language Models

Yutaro Yamada, Yihan Bao, Andrew K. Lampinen, Jungo Kasai, Ilker Yildirim

专题命中 空间理解 :spatial understanding(title)

Comments Accepted to TMLR 2024. Our code and data are available at https://github.com/runopti/SpatialEvalLLM, https://huggingface.co/datasets/yyamada/SpatialEvalLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15800 2023-05-01 cs.HC 71%

UndoPort: Exploring the Influence of Undo-Actions for Locomotion in Virtual Reality on the Efficiency, Spatial Understanding and User Experience

Florian Müller, Arantxa Ye, Dominik Schön, Julian Rasch

专题命中 空间理解 :spatial understanding(title)

Comments To appear in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI 23), April 23-28, 2023, Hamburg, Germany. ACM, New York, NY, USA, 15 pages

Journal ref In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 234, 1-15

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19776 2026-06-19 cs.CV 新提交 70%

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding

Occ-VLM: 面向室内场景理解的占用接地视觉语言模型

Jianing Li, Zhou Fang, Yijiang Liu, Li Du

机构 * School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)

专题命中 空间理解 :3D vision(abstract);point cloud(abstract);分类 cs.CV

AI总结 提出Occ-VLM,仅用姿态RGB图像和单一2D视觉编码器,通过重建3D占用作为几何先验,实现统一的3D场景理解,在占用预测、3D VQA和密集描述任务上达到领先水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05774 2026-06-15 cs.CV 版本更新 70%

LiAuto-GeoX: Efficient Grounded Driving Transformer

LiAuto-GeoX: 高效接地驾驶Transformer

Jiawei Lian, Haoyi Sun, Yang Wu, Lifu Mu, Siyuan Wang, Le Hui, Ning Mao, Tao Wei, Pan Zhou, Kun Zhan, Jian Yang

机构 * Nanjing University of Science and Technology(南京理工大学) Li Auto Inc.(Li Auto公司) Northwestern Polytechnical University(西北工业大学) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学院)

专题命中 空间理解 :3D reconstruction(abstract);spatial understanding(abstract);分类 cs.CV

AI总结 提出LiAuto-GeoX,通过稀疏激光雷达先验和几何保持蒸馏框架,实现高效、实时的自车中心密集3D重建,并显著提升下游自动驾驶任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13394 2026-06-12 cs.RO 新提交 70%

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation

GeoHAT: 几何自适应混合动作Transformer用于移动操作

Xiangyu Zhu, Renjun Wu, Luzhou Ge, Jinyan Liu, Xuesong Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 空间理解 :3D vision(abstract);spatial understanding(abstract);分类 cs.RO

AI总结 提出GeoHAT框架,通过轻量级傅里叶空间编码器注入几何信息,并采用混合全身动作解码器分解机械臂与基座动作,在ManiSkill-HAB基准上成功率提升23.7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02689 2026-04-06 cs.CV cs.AI 70%

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs

Efficient3D: 一种用于3D多模态大语言模型中自适应和去偏token减少的统一框架

Yuhui Lin, Siyue Yu, Yuxing Yang, Guangliang Cheng, Jimin Xiao

机构 * Xi’an Jiaotong-Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学)

专题命中 空间理解 :3D vision(abstract);spatial understanding(abstract);分类 cs.CV

AI总结 本文提出Efficient3D框架,通过去偏视觉token重要性估计器和自适应token再平衡策略,实现3D MLLMs的高效推理,提升性能并降低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24393 2026-03-26 cs.RO 70%

3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models

3D-Mix用于VLA:一种用于将基于VGGT的3D信息整合到视觉-语言-动作模型中的即插即用模块

Bin Yu, Shijie Lian, Xiaopeng Lin, Zhaolong Shen, Yuliang Wei, Haishan Liu, Changti Wu, Hang Yuan, Bailing Wang, Cong Huang, Kai Chen

机构 * HIT(哈尔滨工业大学) ZGCA(中钢集团自动化研究院) ZGCI(中钢集团信息研究院) HUST(华中科技大学) HKUST(GZ)(香港科技大学(广州)) BUAA(北京航空航天大学) ECNU(华东师范大学) DeepCybo

专题命中 空间理解 :3D vision(abstract);spatial understanding(abstract);分类 cs.RO

AI总结 本文提出3D-Mix模块,通过语义条件门控融合机制提升视觉-语言-动作模型的3D感知能力,在标准化基准测试中实现最佳性能。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23607 2026-03-03 cs.CV 70%

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Concerto:联合2D-3D自监督学习产生空间表示

Yujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 空间理解 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV

AI总结 Concerto通过联合2D-3D自监督学习,产生更连贯的信息空间表示,优于现有SOTA模型并在多个基准上取得新成就。

Comments NeurIPS 2025, produced by Pointcept, project page: https://pointcept.github.io/Concerto

Journal ref Neural Information Processing Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17951 2026-02-23 cs.CV cs.AI 70%

ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models

ROCKET:基于残差的多层对齐用于空间感知的视觉-语言-动作模型

Guoheng Sun, Tingting Du, Kaixi Feng, Chenxiang Luo, Xingguo Ding, Zheyu Shen, Ziyao Wang, Yexiao He, Ang Li

机构 * University of Maryland, College Park University of Wisconsin, Madison City University of Hong Kong St.\ Paul's School

专题命中 空间理解 :3D vision(abstract);spatial understanding(abstract);分类 cs.CV

AI总结 ROCKET通过共享投影器和稀疏激活方案实现多层对齐,提升3D空间理解能力,在LIBERO等任务中达到98.5%的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07491 2025-11-06 cs.CV 70%

SpatialLM: Training Large Language Models for Structured Indoor Modeling

Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng, Rui Tang, Hao Zhu, Ping Tan, Zihan Zhou

机构 * Manycore Tech Inc.(Manycore科技公司) Hong Kong University of Science and Technology(香港科技大学)

专题命中 空间理解 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08388 2025-09-11 cs.CV cs.AI 70%

Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

Dubing Chen, Huan Zheng, Yucheng Zhou, Xianfei Li, Wenlong Liao, Tao He, Pai Peng, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学SKL-IOTSC研究所、CIS学院) COWAROBOT Co. Ltd.(COWAROBOT公司)

专题命中 空间理解 :3D vision(abstract);3D reconstruction(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05578 2025-09-09 cs.AI cs.RO 70%

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Ruixun Liu, Lingyu Kong, Derun Li, Hang Zhao

机构 * Shanghai Qi Zhi Institute(上海启智研究所) Xi’an Jiaotong University(西安交通大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 空间理解 :3D vision(abstract);spatial understanding(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03875 2025-04-08 cs.CV 70%

3D Scene Understanding Through Local Random Access Sequence Modeling

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Honglin Chen, Khai Loong Aw, Daniel L. K. Yamins

专题命中 空间理解 :3D vision(abstract);novel view synthesis(abstract);分类 cs.CV

Comments Project webpage: https://neuroailab.github.io/projects/lras_3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00493 2025-03-28 cs.CV cs.CL 70%

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Duo Zheng, Shijia Huang, Liwei Wang

专题命中 空间理解 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01428 2025-03-12 cs.CV 70%

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Zhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang, Hengshuang Zhao

专题命中 空间理解 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV

Comments Project page: https://gpt4scene.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13454 2024-08-27 cs.CV 70%

AdaOcc: Adaptive-Resolution Occupancy Prediction

Chao Chen, Ruoyu Wang, Yuliang Guo, Cheng Zhao, Xinyu Huang, Chen Feng, Liu Ren

专题命中 空间理解 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14215 2024-02-23 cs.CV 70%

Swin3D++: Effective Multi-Source Pretraining for 3D Indoor Scene Understanding

Yu-Qi Yang, Yu-Xiao Guo, Yang Liu

专题命中 空间理解 :3D vision(abstract);point cloud(abstract);分类 cs.CV

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12125 2023-11-22 cs.CV cs.AI 70%

Mixing-Denoising Generalizable Occupancy Networks

Amine Ouasfi, Adnane Boukhayma

专题命中 空间理解 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV

Comments 3DV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10714 2023-05-19 cs.CV 70%

Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

Taolin Zhang, Sunan He, Dai Tao, Bin Chen, Zhi Wang, Shu-Tao Xia

专题命中 空间理解 :3D vision(abstract);point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.05813 2020-11-12 cs.CV 70%

Dynamic Plane Convolutional Occupancy Networks

Stefan Lionar, Daniil Emtsev, Dusan Svilarkovic, Songyou Peng

专题命中 空间理解 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV

Comments To be presented at WACV 2021. Equal contribution between the first three authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.04618 2020-08-04 cs.CV 70%

Convolutional Occupancy Networks

Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, Andreas Geiger

专题命中 空间理解 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV

Comments ECCV 2020 (Spotlight). Project page with supplementary material and code: https://pengsongyou.github.io/conv_onet

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27412 2026-07-21 cs.CV cs.AI cs.GR eess.IV 版本更新 62%

From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation

并非所有关系都同样旋转:面向视角鲁棒的3D场景图生成的变换感知解耦

Jingjun Sun, Chaowei Wang, Zhirui Liu, Jiaxu Tian, Ming Yang, Yaoxing Wang, Yan Di, Shan Gao

机构 * Northwestern Polytechnical University(西北工业大学) ShanghaiTech University(上海科技大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.GR

AI总结 提出变换感知解耦(TAD)框架,通过将关系推理分解为视角稳定和视角变化两部分,结合视角稳定的对象表示,实现3D场景图生成在偏航视角变化下的鲁棒性,无需训练时旋转增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13660 2026-07-07 cs.RO cs.CV 版本更新 62%

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

面向机器人视觉-语言模型的具有推理能力的空间轨迹

Enshen Zhou, Yibo Li, Jingkun An, Jiayuan Zhang, Shanyu Rong, Mengzhen Liu, Yi Han, Yuheng Ji, Huajie Tan, Jiawei He, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Lu Sheng, Shanghang Zhang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 本文提出RoboTracer,一种3D感知的视觉语言模型,通过空间编码器和回归监督解码器提升空间推理能力,并结合强化学习训练实现多步度量推理,最终在复杂场景中实现高效空间轨迹生成。

Comments Accepted to ECCV 2026. Project page: https://zhoues.github.io/RoboTracer

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23565 2026-06-23 cs.RO cs.CV 新提交 62%

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

HoloAgent-0:具有3D空间记忆的统一具身智能体框架

Xiaolin Zhou, Liu Liu, Tingyang Xiao, Wei Feng, Fa Fu, Xinrui Meng, Xinjie Wang, Jialiang Han, Boyang Yu, Yun Du, Wei Sui, Zhizhong Su

机构 * Horizon Robotics D-Robotics Robotics

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 提出HoloAgent-0统一框架,通过三层耦合结构(具身AgentOS、3D空间记忆、具身技能)实现真实机器人闭环执行,在长程导航、跨机器人协调等任务上验证有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17221 2026-06-12 cs.CV cs.RO 版本更新 62%

QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy

QueryOcc:基于查询的3D语义占据自监督方法

Adam Lilja, Ji Lan, Junsheng Fu, Lars Hammarstrand

机构 * Chalmers University of Technology(查尔姆斯理工大学) Zenseact

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 提出QueryOcc,一种基于查询的自监督框架,通过相邻帧的4D时空查询直接学习连续3D语义占据,利用视觉基础模型或激光雷达数据提供监督,并引入收缩场景表示以在恒定内存下实现远程监督,在Occ3D-nuScenes基准上语义RayIoU提升26%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04069 2026-06-02 cs.CV cs.RO 62%

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

SpaceTools: 通过双交互强化学习实现工具增强的空间推理

Siyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak, Adithyavairavan Murali, Qing Qu, Stan Birchfield, Valts Blukis, Jonathan Tremblay

机构 * NVIDIA University of Michigan(密歇根大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 提出双交互强化学习(DIRL)框架,通过两阶段训练让视觉语言模型学会协调多种工具(如深度估计、分割、姿态估计)进行精确空间推理,在多个基准上达到最优性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02060 2026-05-19 cs.CV cs.RO 62%

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

CompassAD: 基于意图的多功能竞争物体3D affordance 地标

Jingliang Li, Jindou Jia, Tuo An, Chuhao Zhou, Xiangyu Chen, Shilin Shan, Boyu Ma, Bofan Lyu, Gen Li, Jianfei Yang

机构 * MARS Lab, Nanyang Technological University, Singapore(MARS实验室,南洋理工大学,新加坡)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 该研究提出了一种新的3D affordance设定,即意图驱动的可混淆地标,旨在预测多物体点云中正确物体的每点affordance掩码,基于隐含的自然语言意图。通过构建CompassAD基准,该研究展示了在具有隐含意图的多物体组合中的先进结果,并在机器人机械臂上验证了其在真实世界抓取中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01365 2026-05-05 cs.CV cs.RO 62%

VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection

VoxAfford:多尺度体素-标记融合用于开放词汇3D affordance检测

Haowen Sun, Shaolong Zhang, Mingyang Li, Chengzhong Ma, Xinzhe Chen, Qiongjie Cui, Xingyu Chen, Zeyang Liu, Xuguang Lan

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家级实验室,人工智能与机器人研究所,西安交通大学)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出VoxAfford,通过多尺度几何特征增强输出标记,提升3D affordance检测的定位精度,实验显示mIoU提升8%,并验证了零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16518 2026-04-29 cs.RO cs.CL cs.CV 62%

MiMo-Embodied: X-Embodied Foundation Model Technical Report

MiMo-Embodied:X-Embodied基础模型技术报告

Xiaoshuai Hao, Lei Zhou, Zhijian Huang, Zhiwen Hou, Yingbo Tang, Lingfeng Zhang, Guang Li, Zheng Lu, Shuhuai Ren, Xianhui Meng, Yuchen Zhang, Jing Wu, Jinghui Lu, Chenxu Dang, Jiayi Guan, Jianhua Wu, Zhiyi Hou, Hanbing Li, Shumeng Xia, Mingliang Zhou, Yinan Zheng, Zihao Yue, Shuhao Gu, Hao Tian, Yuannan Shen, Jianwei Cui, Wen Zhang, Shaoqing Xu, Bing Wang, Haiyang Sun, Zeyu Zhu, Yuncheng Jiang, Zibin Guo, Chuhong Gong, Chaofan Zhang, Wenbo Ding, Kun Ma, Guang Chen, Rui Cai, Diyun Xiang, Heng Qu, Fuli Luo, Hangjun Ye, Long Chen

机构 * Xiaomi Embodied Intelligence Team(小米具身智能团队)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 MiMo-Embodied是首个跨具身基础模型,成功整合并实现自动驾驶和具身AI领域的最先进性能,在17个具身AI基准和12个自动驾驶基准中均取得新纪录,通过多阶段学习、数据构建和CoT/RL微调实现领域间的强正迁移。

Comments Code: https://github.com/XiaomiMiMo/MiMo-Embodied | Model: https://huggingface.co/XiaomiMiMo/MiMo-Embodied-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11236 2026-04-15 cs.CV cs.CL cs.RO 62%

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

ABot-M0:基于动作流形学习的机器人操控VLA基础模型

Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu

机构 * AMAP CV Lab(AMAP视觉实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 ABot-M0通过统一预训练和动作流形学习,提升机器人操控的泛化能力与效率,构建了大规模数据集并支持模块化感知。

Comments Project website: https://amap-cvlab.github.io/ABot-Manipulation/ . Code: https://github.com/amap-cvlab/ABot-Manipulation . 22 pages, 10 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏