arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 7719 信号源:cs.CV, cs.GR, cs.RO

1. 点云 7719 篇

2509.23607 2026-02-18 cs.GR cs.CV 62%

ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing

ZeroScene: 一种基于单张图像的3D场景生成零样本框架及可控纹理编辑

Xiang Tang, Ruotong Li, Xiaopeng Fan

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Pengcheng Laboratory, China(鹏城实验室) Harbin Institute of Technology, China(哈尔滨工业大学) Harbin Institute of Technology, Suzhou Research Institute, China(哈尔滨工业大学苏州市研究院)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.GR

AI总结 ZeroScene通过零样本方法实现单图像3D场景生成与可控纹理编辑,结合大视觉模型先验知识,提升场景连贯性和渲染真实感。

Comments 16 pages, 15 figures, Eurographics 2026, Project page: https://xdlbw.github.io/ZeroScene/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14193 2026-02-17 cs.RO cs.CV cs.LG 62%

Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation

学习具有部分感知的密集3D特征场以实现通用的可变形物体操作

Yue Chen, Muqing Jiang, Kaifeng Zheng, Jiaqi Liang, Chenrui Tie, Haoran Lu, Ruihai Wu, Hao Dong

机构 * Peking University(北京大学) Beijing Institute of Technology(北京理工大学) National University of Singapore(新加坡国立大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出PA3FF,一种具有部分感知的密集3D特征场,用于提升可变形物体操作的泛化能力,通过对比学习训练,实现高效且通用的机器人操作

Comments Accept to ICLR 2026, Project page: https://pa3ff.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08540 2026-02-10 cs.CV cs.GR 62%

TIBR4D: Tracing-Guided Iterative Boundary Refinement for Efficient 4D Gaussian Segmentation

TIBR4D: 基于追踪的迭代边界精修用于高效的4D高斯分割

He Wu, Xia Yan, Yanghui Xu, Liegang Xia, Jiazhou Chen

机构 * Zhejiang University of Technology(浙江工业大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.GR

AI总结 TIBR4D通过两阶段迭代边界精修方法,提升动态4D高斯分割的精度和效率,有效处理遮挡和模糊边界问题。

Comments 13 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03547 2026-02-04 cs.RO cs.CV 62%

AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping

affordanceGrasp-R1: 利用基于推理的 affordance 分割与强化学习进行机器人抓取

Dingyi Zhou, Mu He, Zhuowei Fang, Xiangtong Yao, Yinlong Liu, Alois Knoll, Hu Cao

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 AffordanceGrasp-R1 通过结合推理和强化学习,提升机器人抓取任务中的推理能力和空间定位,实验表明其在基准和现实场景中均表现优异。

Comments Preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00107 2026-02-03 cs.CV cs.RO eess.IV 62%

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

高效无人机轨迹预测:一种多模态深度扩散框架

Yuan Gao, Xinyu Guo, Wenjing Xie, Zifan Wang, Hongwen Yu, Gongyang Li, Shugong Xu

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出一种多模态深度融合框架,通过融合激光雷达和毫米波雷达数据提升无人机轨迹预测精度,实验显示其比基线模型提升40%。

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23107 2026-02-02 cs.CV cs.RO 62%

FlowCalib: LiDAR-to-Vehicle Miscalibration Detection using Scene Flows

FlowCalib: 基于场景流的激光雷达到车辆误校准检测

Ilir Tahiraj, Peter Wittal, Markus Lienkamp

机构 * TUM School of Engineering and Design, Chair of Automotive Technology, Technical University of Munich(慕尼黑技术大学工程与设计学院,汽车技术教授职位) TUM School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算、信息与技术学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 FlowCalib通过分析场景流中的运动线索,首次提出利用神经网络和几何描述符检测激光雷达到车辆的误校准问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01987 2026-02-02 cs.RO cs.CV 62%

CaLiV: LiDAR-to-Vehicle Calibration of Arbitrary Sensor Setups

CaLiV:任意传感器配置的LiDAR到车辆校准

Ilir Tahiraj, Markus Edinger, Dominik Kulmer, Markus Lienkamp

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 CaLiV提出了一种基于目标的校准技术,用于多LiDAR系统的传感器到传感器和传感器到车辆的校准,适用于非重叠视野范围,无需外部设备,实现了高精度的平移和旋转误差校正。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22445 2026-02-02 cs.RO cs.CV 62%

High-Definition 5MP Stereo Vision Sensing for Robotics

高清晰度5MP立体视觉传感用于机器人

Leaf Jiang, Matthew Holzel, Bernhard Kaplan, Hsiou-Yuan Liu, Sabyasachi Paul, Karen Rankin, Piotr Swierczynski

机构 * NODAR Inc.(NODAR公司) NODAR Sensor GmbH(NODAR传感器公司)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本研究提出了一种高精度校准和立体匹配方法,用于提升5MP立体视觉系统在机器人中的应用性能,实现高精度和高速度的3D点云生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09812 2026-01-16 cs.CV cs.RO 62%

LCF3D: A Robust and Real-Time Late-Cascade Fusion Framework for 3D Object Detection in Autonomous Driving

LCF3D: 一种用于自动驾驶中3D物体检测的鲁棒且实时的后期级联融合框架

Carlo Sgaravatti, Riccardo Pieroni, Matteo Corno, Sergio M. Savaresi, Luca Magri, Giacomo Boracchi

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 LCF3D通过结合RGB图像和LiDAR点云的后期级联融合方法,提升自动驾驶中3D物体检测的鲁棒性和实时性。

Comments 35 pages, 14 figures. Published at Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09578 2026-01-15 cs.RO cs.CV 62%

Multimodal Signal Processing For Thermo-Visible-Lidar Fusion In Real-time 3D Semantic Mapping

多模态信号处理用于热-可见-激光雷达融合的实时3D语义制图

Jiajun Sun, Yangyi Ou, Haoyuan Zheng, Chao yang, Yue Ma

机构 * College of Mechatronics and Control Engineering, Shenzhen University(深圳大学机械与控制工程学院) School of Robotics, Xi’an-Jiaotong Liverpool University(西安交通大学利物浦大学机器人学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出通过多模态信号处理融合热、可见和激光雷达数据,提升实时3D语义制图的精度与语义理解能力,适用于灾害评估和工业维护等场景。

Comments 5 pages,7 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09555 2026-01-14 cs.RO cs.CV 62%

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

SpatialActor: 探索解耦的空间表示以实现鲁棒的机器人操作

Hao Shi, Bin Xie, Yingfei Liu, Yang Yue, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang, Gao Huang

机构 * Dexmal

专题命中 点云 :spatial understanding(abstract);分类 cs.CV、cs.RO

AI总结 SpatialActor通过解耦语义和几何,提升机器人操作的鲁棒性和精确性,实现高准确率和强泛化能力。

Comments AAAI 2026 Oral | Project Page: https://shihao1895.github.io/SpatialActor

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15016 2026-01-14 cs.CV cs.RO 62%

MSSF: A 4D Radar and Camera Fusion Framework With Multi-Stage Sampling for 3D Object Detection in Autonomous Driving

MSSF: 一种基于4D雷达和摄像头的多阶段采样融合框架用于自动驾驶中的3D目标检测

Hongsi Liu, Jun Liu, Guangfeng Jiang, Xin Jin

机构 * Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学) Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究所,东部技术研究所)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 MSSF通过多阶段采样融合4D雷达和摄像头数据,提升自动驾驶中3D目标检测的精度与鲁棒性。

Comments T-TITS accepted, code avaliable

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02664 2026-01-12 cs.RO cs.CV cs.LG 62%

Grasp the Graph (GtG) 2.0: Ensemble of Graph Neural Networks for High-Precision Grasp Pose Detection in Clutter

抓取图(GtG)2.0:图神经网络的集合用于在杂乱环境中高精度抓取姿态检测

Ali Rashidi Moghadam, Sayedmohammadreza Rastegari, Mehdi Tale Masouleh, Ahmad Kalhor

机构 * Robot Interaction Laboratory, School of Electrical and Computer Engineering, University of Tehran, Tehran, Iran(机器人交互实验室,电气与计算机工程学院,德黑兰大学,德黑兰,伊朗)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 GtG 2.0通过图神经网络集合提升抓取姿态检测精度,实现高精度和高可靠性

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24384 2026-01-01 cs.RO cs.CV 62%

Geometric Multi-Session Map Merging with Learned Local Descriptors

几何多会话地图融合与学习局部描述符

Yanlong Ma, Nakul S. Joshi, Christa S. Robison, Philip R. Osteen, Brett T. Lopez

机构 * University of California, Los Angeles(加州大学洛杉矶分校) DEVCOM Army Research Laboratory (ARL)(国防部陆军研究实验室(ARL))

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 本文提出GMLD框架,通过学习局部描述符实现大规模多会话点云地图融合,提升地图对齐的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02903 2025-12-30 cs.CV cs.RO 62%

LidarDM: Generative LiDAR Simulation in a Generated World

LidarDM: 生成世界中的生成激光雷达模拟

Vlas Zyrianov, Henry Che, Zhijian Liu, Shenlong Wang

机构 * UIUC(伊利诺伊大学香槟分校) NVIDIA(英伟达)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 LidarDM通过集成4D世界生成框架,实现了基于驾驶场景的4D激光雷达点云生成,提升自动驾驶模拟的逼真度与时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16724 2025-12-19 cs.RO cs.CV 62%

VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation

VERM:利用基础模型创建虚拟眼睛以实现高效的3D机器人操作

Yixiang Chen, Yan Huang, Keji He, Peiyan Li, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新技术实验室,多模态人工智能系统国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges Shandong University(山东大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 VERM通过利用基础模型创建虚拟视图,提升3D机器人操作的效率和准确性,实现训练和推理速度的显著提升。

Comments Accepted at RA-L 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16349 2025-12-17 cs.CV cs.GR 62%

CRISTAL: Real-time Camera Registration in Static LiDAR Scans using Neural Rendering

CRISTAL:利用神经渲染进行静态LiDAR扫描中的实时相机定位

Joni Vanherck, Steven Moonen, Brent Zoomers, Kobe Werner, Jeroen Put, Lode Jorissen, Nick Michiels

机构 * Hasselt University - Digital Future Lab - Flanders(哈塞尔特大学-数字未来实验室-弗拉芒)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.GR

AI总结 CRISTAL通过神经渲染技术实现静态LiDAR扫描中的实时相机定位,提供无漂移且具有正确度量尺度的跟踪,优于现有SLAM流程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19912 2025-12-09 cs.CV cs.LG cs.RO 62%

Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data Pretraining

增强的时空一致性用于图像到LiDAR数据预训练

Xiang Xu, Lingdong Kong, Hui Shuai, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu, Qingshan Liu

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Computing, Department of Computer Science, National University of Singapore(新加坡国立大学计算机学院) School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机学院) Shanghai AI Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 SuperFlow++通过整合时空线索提升图像到LiDAR数据预训练效果,实现更鲁棒的特征表示和更高效的自动驾驶感知。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04005 2025-12-04 cs.CV cs.LG cs.RO 62%

LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving

LargeAD: 大规模跨传感器数据预训练用于自动驾驶

Lingdong Kong, Xiang Xu, Youquan Liu, Jun Cen, Runnan Chen, Wenwei Zhang, Liang Pan, Kai Chen, Ziwei Liu

机构 * WorldBench Team(WorldBench团队)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 LargeAD通过跨传感器数据预训练提升自动驾驶中的三维场景理解,结合多模态对比学习和时间一致性,实现更鲁棒的感知性能。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02972 2025-12-03 cs.CV cs.RO 62%

BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection

BEVDilation:以LiDAR为中心的多模态融合用于3D目标检测

Guowen Zhang, Chenhang He, Liyi Chen, Lei Zhang

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 BEVDilation提出以LiDAR为中心的多模态融合方法,通过稀疏体素扩张和语义引导BEV扩张模块提升3D目标检测性能,有效缓解深度误差带来的空间错位问题。

Comments Accept by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22171 2025-12-01 cs.CV cs.GR 62%

BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch

BrepGPT: 基于Voronoi半块的自回归B-rep生成

Pu Li, Wenhao Zhang, Weize Quan, Biao Zhang, Peter Wonka, Dong-Ming Yan

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) University of Chinese Academy of Sciences(中国科学院大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.GR

AI总结 BrepGPT通过Voronoi半块表示实现单阶段自回归B-rep生成,提升生成效率与模型紧凑性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19684 2025-11-26 cs.CV cs.AI cs.HC cs.RO 62%

IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants

IndEgo:工业场景及协作工作的人际助理数据集

Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai, Bosong Wang, Ze Lu, Oliver Heimann, Jörg Krüger

机构 * Fraunhofer IPK(弗劳恩霍夫研究所) Technical University of Berlin(柏林技术大学) University of Tübingen(图宾根大学) RWTH Aachen University(亚琛工业大学) Leibniz University Hannover(汉诺威莱布尼茨大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 IndEgo数据集旨在通过工业场景中的协作工作提升人机交互助理的多模态理解与错误检测能力。

Comments Accepted to NeurIPS 2025 D&B Track. Project Page: https://indego-dataset.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17225 2025-11-24 cs.RO cs.AI cs.CV 62%

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

TP-MDDN:任务优先的多需求驱动导航与自主决策

Shanshan Li, Da Huang, Yu He, Yanwei Fu, Yu-Gang Jiang, Xiangyang Xue

机构 * Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institution(上海创新机构)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

AI总结 TP-MDDN通过引入任务优先的多需求驱动导航与自主决策系统,提升复杂任务下的导航性能与环境理解能力。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01196 2025-11-19 cs.RO cs.AI cs.CV 62%

OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model

Ishika Singh, Ankit Goyal, Stan Birchfield, Dieter Fox, Animesh Garg, Valts Blukis

机构 * University of Southern California(美国南加州大学) NVIDIA(英伟达)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22344 2025-11-18 cs.CV cs.RO 62%

Task-Driven Implicit Representations for Automated Design of LiDAR Systems

Nikhil Behari, Aaron Young, Tzofi Klinghoffer, Akshat Dave, Ramesh Raskar

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 点云 :3D vision(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06378 2025-11-11 cs.RO cs.CV 62%

ArtReg: Visuo-Tactile based Pose Tracking and Manipulation of Unseen Articulated Objects

Prajval Kumar Murali, Mohsen Kaboli

机构 * RoboTac Lab, BMW Group(RoboTac实验室,宝马集团) University of Glasgow(格拉斯哥大学) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12310 2025-11-11 cs.CV cs.AI cs.RO 62%

DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations

Shouyi Lu, Huanyu Zhou, Guirong Zhuo, Xiao Tang

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

Comments 9 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00801 2025-11-11 cs.CV cs.AI cs.RO 62%

Environment-Driven Online LiDAR-Camera Extrinsic Calibration

Zhiwei Huang, Jiaqi Li, Hongbo Zhao, Xiao Ma, Ping Zhong, Xiaohu Zhou, Wei Ye, Rui Fan

机构 * Department of Control Science & Engineering, the College of Electronic & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) Beijing Institute of Aerospace Control Devices(北京航天控制器件研究所) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15530 2025-11-04 cs.RO cs.CV cs.LG 62%

VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation

Zehao Ni, Yonghao He, Lingfeng Qian, Jilei Mao, Fa Fu, Wei Sui, Hu Su, Junran Peng, Zhipeng Wang, Bin He

机构 * D-Robotics(D机器人) National Key Laboratory of Autonomous Intelligent Unmanned Systems(国家级自主智能无人机系统实验室) University of Science and Technology Beijing(北京科技大学) State Key Laboratory of Multimodal Artificial Intelligence System (MAIS) Institute of Automation of Chinese Academy of Sciences(多模态人工智能系统(MAIS)实验室,中国科学院自动化研究所) Frontiers Science Center for Intelligent Autonomous Systems(智能自主系统前沿科学中心) Shanghai Institute of Intelligent Science and Technology, Tongji University(上海智能科学技术研究院,同济大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00060 2025-11-04 cs.CV cs.RO eess.IV 62%

Which LiDAR scanning pattern is better for roadside perception: Repetitive or Non-repetitive?

Zhiqi Qi, Runxin Zhao, Hanyang Zhuang, Chunxiang Wang, Ming Yang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) Global College, Shanghai Jiao Tong University(全球学院,上海交通大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏