arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 497 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 497 篇

2505.05800 2026-03-31 cs.RO cs.CV 62%

3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks

3D CAVLA:利用深度和3D上下文来泛化视觉语言动作模型以应对未见任务

Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

机构 * New York University Tandon School of Engineering(纽约大学坦登工程学院) New York University Courant Institute of Mathematical Sciences(纽约大学库朗数学科学研究所)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 本文提出3D-CAVLA框架,通过引入链式推理、深度感知和任务导向的感兴趣区域检测,提升视觉语言动作模型在未见任务中的泛化能力,实验表明其在模拟和现实任务中均表现出色。

Comments Accepted at the 1st Workshop on 3D LLM/VLA, CVPR 2025. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14498 2026-03-31 cs.RO cs.CV 62%

R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation

R3DP:实时3D感知策略用于具身操作

Yuhao Zhang, Wanxi Dong, Yue Shi, Yi Liang, Jingnan Gao, Qiaochu Yang, Yaxing Lyu, Zhixuan Liang, Yibin Liu, Congsheng Xu, Xianda Guo, Wei Sui, Yaohui Jin, Xiaokang Yang, Yanyan Xu, Yao Mu

机构 * Shanghai Jiao Tong University(上海交通大学) D-Robotics(地平线机器人) Southern University of Science and Technology(南方科技大学) Xspark AI(星火科技) The University of Hong Kong(香港大学) Wuhan University(武汉大学) Xiamen University Malaysia(厦门大学马来西亚分校) Northeastern University(东北大学)

专题命中 空间理解 :3D vision(abstract);分类 cs.CV、cs.RO

AI总结 R3DP通过异步快慢协作模块整合大尺度3D先验知识,提升实时操作性能,实验显示在成功率和推理时间上均优于现有方法。

Comments Project Page: https://dazazh.github.io/r3dp-project-page/ Github Repo: https://github.com/dazazh/R3DP

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24581 2026-03-26 cs.CV cs.RO 62%

Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving

潜在世界动作建模:端到端自动驾驶的潜在世界建模

Linbo Wang, Yupeng Zheng, Qiang Chen, Shiwei Li, Yichen Zhang, Zebin Xing, Qichao Zhang, Xiang Li, Deheng Qian, Pengxuan Yang, Yihang Dong, Ce Hao, Xiaoqing Ye, Junyu han, Yifeng Pan, Dongbin Zhao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Chongqing Chang’an Technology Co., Ltd(重庆长安科技有限公司) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) College of AI, Tsinghua University(清华大学人工智能学院) Zhongguancun Academy(中关村学院)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 本文提出Latent-WAM框架,通过空间感知和动态感知的潜在世界表示实现高效端到端自动驾驶,实验显示在NAVSIM v2和HUGSIM上取得新的SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22280 2026-03-24 cs.CV cs.RO 62%

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

DualCoT-VLA: 通过并行推理实现视觉-语言链式思维的视觉-语言-动作模型

Zhide Zhong, Junfeng Li, Junjie He, Haodong Yan, Xin Gong, Guanyi Zhao, Yingjie Cai, Jiantao Gao, Xu Yan, Bingbing Liu, Yingcong Chen, Liuqing Yang, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Huawei Foundation Model Department(华为基础模型部门)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 DualCoT-VLA通过并行推理机制,结合视觉和语言链式思维,解决传统VLA模型在复杂多步骤任务和精细空间感知中的不足,实现更高效的视觉-语言-动作处理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02609 2026-03-04 cs.CV cs.RO 62%

VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semantic Occupancy Prediction

VLMFusionOcc3D: 基于VLM的多模态3D语义占位预测

A. Enes Doruk, Hasan F. Ates

机构 * Department of Artificial Intelligence and Data Engineering, Ozyegin University(人工智能与数据工程系,奥克辛大学)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 VLMFusionOcc3D通过融合视觉语言模型的语义先验,提升多模态3D语义占位预测的鲁棒性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00551 2026-02-03 cs.RO cs.CV 62%

APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation

APEX:一种解耦的记忆型探索者用于异步空中目标导航

Daoxuan Zhang, Ping Chen, Xiaobo Xia, Xiu Su, Ruichen Zhen, Jianqiang Xiao, Shuo Yang

机构 * Harbin Institute of Technology(哈尔滨工业大学) National University of Singapore(新加坡国立大学) Central South University(中南大学) Meituan Academy of Robotics(美团机器人研究院)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 APEX通过分层异步架构实现高效空中目标导航,提升探索效率与目标识别精度。

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17885 2026-01-27 cs.CV cs.AI cs.RO 62%

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

PEAfowl:感知增强的多视角视觉-语言-动作用于双臂操作

Qingyu Fan, Zhaoxiang Li, Yi Lu, Wang Chen, Qiu Shen, Xiao-xiao Long, Yinghao Cai, Tao Lu, Shuo Wang, Xun Cao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Nanjing University(南京大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 PEAfowl通过增强感知的多视角视觉-语言-动作策略,提升双臂操作在复杂环境中的稳定性和成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16262 2026-01-06 cs.RO cs.CV 62%

How Robot Dogs See the Unseeable: Improving Visual Interpretability via Peering for Exploratory Robots

机器人如何看见不可见之物:通过窥视提升探索机器人视觉可解释性

Oliver Bimber, Karl Dietrich von Ellenrieder, Michael Haller, Rakesh John Amala Arokia Nathan, Gianni Lunardi, Mohamed Youssef, Marco Camurri, Santos Miguel Orozco Soto, Jeremy E. Niven

专题命中 空间理解 :3D vision(abstract);分类 cs.CV、cs.RO

AI总结 通过模仿昆虫的窥视动作,提升探索机器人在部分遮挡下的视觉推理能力,实现高分辨率、实时感知,适用于复杂环境导航与场景理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04308 2026-01-06 cs.RO cs.AI cs.CV 62%

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

RoboRefer: 向视觉语言模型在机器人中的空间指称推理迈进

Enshen Zhou, Jingkun An, Cheng Chi, Yi Han, Shanyu Rong, Chi Zhang, Pengwei Wang, Zhongyuan Wang, Tiejun Huang, Lu Sheng, Shanghang Zhang

机构 * School of Software, Beihang University(北京航空航天大学软件学院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 RoboRefer通过整合深度编码器和强化微调方法,实现了视觉语言模型在机器人中的空间指称推理,提升了复杂场景下的交互能力。

Comments Accepted by NeurIPS 2025. Project page: https://zhoues.github.io/RoboRefer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13142 2026-01-01 cs.CV cs.CL cs.LG cs.MM cs.RO 62%

Holistic Evaluation of Multimodal LLMs on Spatial Intelligence

多模态大语言模型在空间智能方面的综合评估

Zhongang Cai, Yubo Wang, Qingping Sun, Ruisi Wang, Chenyang Gu, Wanqi Yin, Zhiqian Lin, Zhitao Yang, Chen Wei, Oscar Qian, Hui En Pang, Xuanke Shi, Kewang Deng, Xiaoyang Han, Zukai Chen, Jiaqi Li, Xiangyu Fan, Hanming Deng, Lewei Lu, Bo Li, Ziwei Liu, Quan Wang, Dahua Lin, Lei Yang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 本文提出EASI框架,用于评估多模态大语言模型在空间智能方面的综合表现,揭示GPT-5在SI中的优势与不足,并展示专有模型在困难任务上的劣势。

Comments Codebase: https://github.com/EvolvingLMMs-Lab/EASI/ ; Leaderboard: https://huggingface.co/spaces/lmms-lab-si/EASI-Leaderboard

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16461 2025-12-19 cs.CV cs.RO 62%

SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning

SNOW:基于世界知识的时空场景理解用于开放世界具身推理

Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Esslingen University of Applied Sciences(埃斯林根应用科学大学) Dr. Ing. h.c. F. Porsche AG(德意志联邦汽车工业联合会) University of Michigan(密歇根大学) Voxel51 Inc.(Voxel51公司)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 SNOW通过整合视觉语言模型与点云几何,实现统一的4D场景理解,提升开放世界具身推理的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13177 2025-12-17 cs.CV cs.RO 62%

MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

MMDrive: 通过多表示融合超越视觉的交互场景理解

Minghui Hou, Wei-Hsing Huang, Shaofeng Liang, Daizong Liu, Tai-Hao Wen, Gang Wang, Runwei Guan, Weiping Ding

机构 * organization= College of Computer Science Technology, Jilin University , city= Changchun , country= China organization= Georgia Institute of Technology , city= Atlanta , country= USA organization= Qingdao Institute of Software, College of Computer Science Technology, China University of Petroleum (East China) , city= Qingdao , country= China organization= Institute for Math \& AI, Wuhan University , city= Wuhan , country= China organization= University of Michigan, Ann Arbor , country= USA organization= Thrust of Artificial Intelligence, Hong Kong University of Science organization= School of Artificial Intelligence Computer Science, Nantong University , city= Nantong , country= China

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 MMDrive通过融合占用图、LiDAR点云和文本描述,实现超越视觉的三维场景理解,提升自动驾驶的多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05258 2025-12-08 cs.CV cs.LG cs.RO 62%

Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving

多模态数据高效3D场景理解用于自动驾驶

Lingdong Kong, Xiang Xu, Jiawei Ren, Wenwei Zhang, Liang Pan, Kai Chen, Wei Tsang Ooi, Ziwei Liu

机构 * WorldBench Team Project Lead(WorldBench团队项目负责人)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 LaserMix++通过多模态方法提升自动驾驶中LiDAR数据高效3D场景理解,以更少标注实现更高精度。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23186 2025-12-01 cs.RO cs.AI cs.CV 62%

Obstruction reasoning for robotic grasping

机器人抓取中的障碍物推理

Runyu Jiao, Matteo Bortolon, Francesco Giuliari, Alice Fasoli, Sergio Povoli, Guofeng Mei, Yiming Wang, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布鲁诺·科塞勒基金会) University of Trento(特伦托大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 UNOGrasp通过视觉-语言模型提升机器人抓取中障碍物推理能力,结合监督与强化学习微调,实现更高效的路径规划和抓取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05622 2025-11-11 cs.CV cs.AI cs.LG cs.RO 62%

Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition

Nicholas Babey, Tiffany Gu, Yiheng Li, Cristian Meo, Kevin Zhu

机构 * Arizona State University(亚利桑那州立大学) Emory University(埃默里大学) University of California, Berkeley(加州大学伯克利分校) Delft University of Technology(代尔夫特理工大学) Algoverse AI Research(Algoverse AI 研究)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments Accepted at NeurIPS 2025 SpaVLE, for code see https://github.com/nbabey20/groundactrec , 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02348 2025-10-15 cs.CV cs.AI cs.RO 62%

mmWave Radar-Based Non-Line-of-Sight Pedestrian Localization at T-Junctions Utilizing Road Layout Extraction via Camera

Byeonggyu Park, Hee-Yeun Kim, Byonghyok Choi, Hansang Cho, Byungkwan Kim, Soomok Lee, Mingu Jeon, Seong-Woo Kim

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16721 2025-09-23 cs.CV cs.AI cs.RO 62%

Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding

Haoyuan Li, Rui Liu, Hehe Fan, Yi Yang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 空间理解 :3D vision(abstract);分类 cs.CV、cs.RO

Comments 19 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10880 2025-09-23 cs.CV cs.RO 62%

Safe-Construct: Redefining Construction Safety Violation Recognition as 3D Multi-View Engagement Task

Aviral Chharia, Tianyu Ren, Tomotake Furuhata, Kenji Shimada

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments CVPR Workshop 2025; Project Website: https://Safe-Construct.github.io/Safe-Construct

Journal ref CVPR, Nashville, TN, USA, 2025, pp. 5811-5820

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09594 2025-09-12 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

ObjectReact: Learning Object-Relative Control for Visual Navigation

Sourav Garg, Dustin Craggs, Vineeth Bhat, Lachlan Mares, Stefan Podgorski, Madhava Krishna, Feras Dayoub, Ian Reid

机构 * University of Adelaide, Australia(澳大利亚阿德莱德大学) IIIT Hyderabad, India(印度海得拉巴IIIT) MBZUAI, UAE(阿联酋MBZUAI)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments CoRL 2025; 23 pages including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01106 2025-09-12 cs.AI cs.CV cs.RO 62%

Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Huang Fang, Mengxi Zhang, Heng Dong, Wei Li, Zixuan Wang, Qifeng Zhang, Xueyun Tian, Yucheng Hu, Hang Li

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments Tech report. Project page: https://robix-seed.github.io/robix/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08302 2025-09-11 cs.RO cs.CV 62%

Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

Rajendramayavan Sathyam, Yueqi Li

机构 * Zoox Inc.(Zoox公司)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments 32 pages, 14 figures, accepted at IEEE Open Journal of Vehicular Technology (OJVT)

URL PDF HTML 收藏
2505.12363 2025-09-10 cs.CV cs.AI cs.CL cs.LG cs.RO 62%

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

Qi Feng

机构 * Kyoto University(京都大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

Comments 26 pages, 19 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06233 2025-09-09 cs.RO cs.CV 62%

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

Tongxuan Tian, Xuhui Kang, Yen-Ling Kuo

机构 * University of Virginia(弗吉尼亚大学)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

Comments Conference on Robot Learning (CoRL) 2025. Project website: https://o3afford.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01197 2025-09-04 cs.CV cs.RO 62%

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu

机构 * College of Software Technology, Zhejiang University(浙江大学软件技术学院) College of Computer Science, Zhejiang University(浙江大学计算机科学学院)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15472 2025-08-05 cs.CV cs.AI cs.GR 62%

KinMo: Kinematic-aware Human Motion Understanding and Generation

Pengfei Zhang, Pinxin Liu, Pablo Garrido, Hyeongwoo Kim, Bindita Chaudhuri

机构 * University of California, Irvine(加州大学欧文分校) University of Rochester(罗切斯特大学) Imperial College, London(伦敦帝国学院) Flawless AI

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.GR

Comments Accepted to ICCV 2025; Project page: https://andypinxinliu.github.io/KinMo

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13679 2025-06-17 cs.RO cs.AI cs.CV 62%

ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Yuqing Wen, Kefan Gu, Haoxuan Liu, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学) Dexmal

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17735 2025-04-07 cs.CV cs.RO 62%

3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning

Yuncong Yang, Han Yang, Jiachen Zhou, Peihao Chen, Hongxin Zhang, Yilun Du, Chuang Gan

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12663 2025-03-18 cs.CV cs.CL cs.LG cs.RO 62%

Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding

Imran Kabir, Md Alimoor Reza, Syed Billah

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01594 2024-12-17 cs.CV cs.GR cs.LG 62%

DiffUHaul: A Training-Free Method for Object Dragging in Images

Omri Avrahami, Rinon Gal, Gal Chechik, Ohad Fried, Dani Lischinski, Arash Vahdat, Weili Nie

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.GR

Comments Accepted to SIGGRAPH Asia 2024. Project page is available at https://omriavrahami.com/diffuhaul/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02449 2024-12-04 cs.RO cs.AI cs.CV cs.LG 62%

BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding

Chenguang Huang, Shengchao Yan, Wolfram Burgard

专题命中 空间理解 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏