arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

2026-03-31 至 2026-03-31 共收录 4 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 4 篇

2505.05800 2026-03-31 cs.RO cs.CV 62%

3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks

3D CAVLA:利用深度和3D上下文来泛化视觉语言动作模型以应对未见任务

Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

机构 * New York University Tandon School of Engineering(纽约大学坦登工程学院) New York University Courant Institute of Mathematical Sciences(纽约大学库朗数学科学研究所)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 本文提出3D-CAVLA框架,通过引入链式推理、深度感知和任务导向的感兴趣区域检测,提升视觉语言动作模型在未见任务中的泛化能力,实验表明其在模拟和现实任务中均表现出色。

Comments Accepted at the 1st Workshop on 3D LLM/VLA, CVPR 2025. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14498 2026-03-31 cs.RO cs.CV 62%

R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation

R3DP:实时3D感知策略用于具身操作

Yuhao Zhang, Wanxi Dong, Yue Shi, Yi Liang, Jingnan Gao, Qiaochu Yang, Yaxing Lyu, Zhixuan Liang, Yibin Liu, Congsheng Xu, Xianda Guo, Wei Sui, Yaohui Jin, Xiaokang Yang, Yanyan Xu, Yao Mu

机构 * Shanghai Jiao Tong University(上海交通大学) D-Robotics(地平线机器人) Southern University of Science and Technology(南方科技大学) Xspark AI(星火科技) The University of Hong Kong(香港大学) Wuhan University(武汉大学) Xiamen University Malaysia(厦门大学马来西亚分校) Northeastern University(东北大学)

专题命中 空间理解 :3D vision(abstract);分类 cs.CV、cs.RO

AI总结 R3DP通过异步快慢协作模块整合大尺度3D先验知识,提升实时操作性能,实验显示在成功率和推理时间上均优于现有方法。

Comments Project Page: https://dazazh.github.io/r3dp-project-page/ Github Repo: https://github.com/dazazh/R3DP

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27970 2026-03-31 cs.CV 57%

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

AffordMatcher: 从视觉符号中学习3D场景中的 affordance

Nghia Vu, Tuong Do, Khang Nguyen, Baoru Huang, Nhat Le, Binh Xuan Nguyen, Erman Tjiputra, Quang D. Tran, Ravi Prakash, Te-Chuan Chiu, Anh Nguyen

机构 * University of Liverpool(利物浦大学) AIOZ Ltd.(AIOZ有限公司) National Tsing Hua University(国立清华大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学) University of Western Australia(西澳大学) Indian Institute of Science(印度科学理工学院) NVIDIA(英伟达)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 本文提出AffordMatcher,通过结合点云和图像实例,利用视觉符号实现更精确的affordance区域识别,基于大规模数据集验证了方法有效性。

Comments 14 pages. Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27294 2026-03-31 cs.CV 57%

Class-Distribution Guided Active Learning for 3D Occupancy Prediction in Autonomous Driving

基于类分布的主动学习用于自动驾驶中的3D占用预测

Wonjune Kim, In-Jae Lee, Sihwan Hwang, Sanmin Kim, Dongsuk Kum

机构 * KAIST(韩国科学技术院) Seoul National University(首尔大学) Kookmin University(国民大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出一种基于类分布的主动学习框架,通过三种互补标准选择训练样本,以解决自动驾驶中3D占用预测的类别不平衡和标注成本问题,实现在仅42.4%标注数据下达到26.62 mIoU的性能。

Comments IEEE RA-L 2026

Journal ref IEEE Robotics and Automation Letters (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏