arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

2026-04-17 至 2026-04-17 共收录 16 信号源:cs.CV, cs.GR, cs.RO

1. 三维重建 1 篇

2604.14141 2026-04-17 cs.CV 83%

Geometric Context Transformer for Streaming 3D Reconstruction

流式3D重建的几何上下文变换器

Lin-Zhuo Chen, Jian Gao, Yihang Chen, Ka Leong Cheng, Yipengjing Sun, Liangxiao Hu, Nan Xue, Xing Zhu, Yujun Shen, Yao Yao, Yinghao Xu

机构 * technology.robbyant.com(robbyant技术公司)

专题命中 三维重建 :3D reconstruction(title,abstract);point cloud(abstract);分类 cs.CV

AI总结 本文提出LingBot-Map,基于几何上下文变换器架构,通过精心设计的注意力机制实现流式3D重建,兼顾几何精度、时间一致性和计算效率,达到20FPS的稳定高效推理。

Comments Project page: https://technology.robbyant.com/lingbot-map Code: https://github.com/robbyant/lingbot-map

详情

展开后加载摘要…

URL PDF HTML 收藏

2. Gaussian Splatting 5 篇

2604.14706 2026-04-17 cs.CV 94%

NG-GS: NeRF-Guided 3D Gaussian Splatting Segmentation

NG-GS: 基于NeRF引导的3D高斯点云分割

Yi He, Tao Wang, Yi Jin, Congyan Lang, Yidong Li, Haibin Ling

机构 * Key Laboratory of Big Data and Artificial Intelligence in Transportation, Ministry of Education(交通大数据与人工智能联合实验室,教育部) School of Computer and Information Technology, Beijing Jiaotong University(北京交通大学计算机与信息学院) Department of Artificial Intelligence, Westlake University(西湖大学人工智能学院)

专题命中 Gaussian Splatting :NeRF(title,title_cn);Gaussian Splatting(title,abstract);3DGS(abstract,abstract_cn);novel view synthesis(abstract)

AI总结 本文提出NG-GS框架,通过掩码方差分析识别边界模糊的高斯点,利用RBF插值和多分辨率哈希编码构建连续特征场,结合NeRF模块实现高质量3D高斯点云分割。

Comments Accepted to CVPR 2026 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12580 2026-04-17 cs.CV 85%

PDF-GS: Progressive Distractor Filtering for Robust 3D Gaussian Splatting

PDF-GS:渐进式干扰过滤用于鲁棒的3D高斯点划法

Kangmin Seo, MinKyu Lee, Tae-Young Kim, ByeongCheol Lee, JoonSeoung An, Jae-Pil Heo

机构 * Sungkyunkwan University(成均馆大学)

专题命中 Gaussian Splatting :Gaussian Splatting(title,abstract);3DGS(abstract,abstract_cn);分类 cs.CV

AI总结 PDF-GS通过渐进多阶段优化增强3D高斯点划法的自过滤能力,有效去除干扰,实现高质量重建,优于现有方法。

Comments Accepted to CVPR Findings 2026. Project Page: https://kangrnin.github.io/PDF-GS

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15239 2026-04-17 cs.CV 84%

TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens

TokenGS: 通过可学习的标记解耦3D高斯预测与像素

Jiawei Ren, Michal Jan Tyszkiewicz, Jiahui Huang, Zan Gojcic

机构 * NVIDIA

专题命中 Gaussian Splatting :3DGS(summary_cn,abstract);Gaussian Splatting(abstract);分类 cs.CV

AI总结 TokenGS通过自监督渲染损失直接回归3D均值坐标,改进了3DGS预测的鲁棒性与效率,实现了更规则的几何和更平衡的3DGS分布,支持多视图不一致性和姿态噪声的鲁棒性。

Comments Project page: https://research.nvidia.com/labs/toronto-ai/tokengs

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14268 2026-04-17 cs.CV 81%

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

HY-World 2.0:一个多模态世界模型用于重建、生成和模拟3D世界

Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, Bo Yuan, Chao Zhang, Coopers Li, Dongyuan Guo, Fan Yang, Haiyu Zhang, Hang Cao, Jianchen Zhu, Jiaxin Lin, Jie Xiao, Jihong Zhang, Junlin Yu, Lei Wang, Lifu Wang, Lilin Wang, Linus, Minghui Chen, Peng He, Penghao Zhao, Qi Chen, Rui Chen, Rui Shao, Sicong Liu, Wangchen Qin, Xiaochuan Niu, Xiang Yuan, Yi Sun, Yifei Tang, Yifu Sun, Yihang Lian, Yonghao Tan, Yuhong Liu, Yuyang Yin, Zhiyuan Min, Tengfei Wang, Chunchao Guo

机构 * Tencent(腾讯)

专题命中 Gaussian Splatting :Gaussian Splatting(abstract,abstract_cn);3DGS(abstract,abstract_cn);分类 cs.CV

AI总结 HY-World 2.0通过多模态输入生成高保真3D场景,改进了先前版本,引入了多项创新以提升全景真实性、3D场景理解和规划能力,并升级了WorldStereo和WorldMirror模型,实现了3D世界交互探索。

Comments Project Page: https://3d-models.hunyuan.tencent.com/world/ ; Code: https://github.com/Tencent-Hunyuan/HY-World-2.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14782 2026-04-17 cs.CV 77%

One-shot Compositional 3D Head Avatars with Deformable Hair

单张图像生成可组合的3D头部化身与可变形头发

Yuan Sun, Xuan Wang, WeiLi Zhang, Wenxuan Zhang, Yu Guo, Fei Wang

机构 * Xi'an Jiaotong University(西安交通大学)

专题命中 Gaussian Splatting :3DGS(abstract,abstract_cn);Gaussian Splatting(abstract);分类 cs.CV

AI总结 本文提出一种单张图像生成完整3D头部化身的方法,通过分离头发与面部区域,结合图像到3D提升技术,实现更逼真的头发动态和面部细节。

Comments project page: https://yuansun-xjtu.github.io/CompHairHead.io

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 点云 6 篇

2604.15023 2026-04-17 cs.RO 57%

DockAnywhere: Data-Efficient Visuomotor Policy Learning for Mobile Manipulation via Novel Demonstration Generation

DockAnywhere: 通过新颖的示范生成实现移动操作的数据高效视觉-运动策略学习

Ziyu Shan, Yuheng Zhou, Gaoyuan Wu, Ziheng Ji, Zhenyu Wu, Ziwei Wang

机构 * Nanyang Technological University, Singapore(新加坡南洋理工大学) Beijing University of Posts and Telecommunications, Beijing, China(北京邮电大学)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出DockAnywhere方法,通过解耦基座运动与不变的 manipulation 技能,提升移动操作在不同 docking 点下的视角泛化能力,实验表明其显著提高了策略成功率和泛化能力。

Comments Accepted to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14781 2026-04-17 cs.CV 57%

Integrating Object Detection, LiDAR-Enhanced Depth Estimation, and Segmentation Models for Railway Environments

集成目标检测、LiDAR增强的深度估计和分割模型用于铁路环境

Enrico Francesco Giannico, Federico Nesti, Gianluca D'Amico, Mauro Marinoni, Edoardo Carosio, Filippo Salotti, Salvatore Sabina, Giorgio Buttazzo

机构 * 1 Department of Excellence in Robotics \& AI, Scuola Superiore Sant'Anna, Pisa, Italy 2 Progress Rail Signaling S.p.A., Serravalle Pistoiese (PT), Italy

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出一个模块化框架,结合目标检测、轨道分割和单目深度估计,用于铁路环境中的障碍物检测与距离估计,通过合成数据集评估,实现0.63米的高精度距离估计。

Comments Under submission for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14635 2026-04-17 cs.RO 57%

A multi-platform LiDAR dataset for standardized forest inventory measurement at long term ecological monitoring sites

多平台激光雷达数据集用于长期生态监测站点标准化森林调查测量

Michael R. Chang, Anna Candotti, Karl von Ellenrieder, Enrico Tomelleri, Marco Camurri

机构 * Faculty of Engineering, Free University of Bozen-Bolzano(博尔扎诺自由大学工程学院) Department of Industrial Engineering, University of Trento(特伦托大学工业工程系) Faculty of Agricultural, Environmental and Food Sciences, Free University of Bozen-Bolzano(博尔扎诺自由大学农业、环境与食品科学学院) Competence Centre for Mountain Innovation Ecosystems, Free University of Bozen-Bolzano(博尔扎诺自由大学山地创新生态系统能力中心)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出一个多平台激光雷达数据集,用于校准、基准测试和整合3D结构数据与生态观测及标准生物量模型,结合无人机、地面和背包式激光扫描技术,提供高精度森林结构数据。

Comments 30 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03093 2026-04-17 cs.CV 57%

Estimating the Diameter at Breast Height of Trees in a Forest from RGB

从RGB估算森林中树木的胸径

Siming He, Zachary Osman, Fernando Cladera, Dexter Ong, Nitant Rai, Patrick Corey Green, Vijay Kumar, Pratik Chaudhari

机构 * General Robotics, Automation, Sensing and Perception (GRASP) Laboratory, University of Pennsylvania(通用机器人、自动化、传感与感知实验室,宾夕法尼亚大学) Department of Forest Resources and Environmental Conservation, Virginia Tech(森林资源与环境保护系,弗吉尼亚理工学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出一种低成本方法,利用消费级360视频相机估算树木胸径,结合SfM摄影测量和语义分割技术,实现高精度测量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24886 2026-04-17 cs.LG 50%

Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks

自适应规范化解析及其在不变各向异性几何网络中的应用

Ya-Wei Eileen Lin, Ron Levie

机构 * Technical University of Munich, School of Computation, Information and Technology(慕尼黑技术大学,计算、信息与技术学院) Munich Center for Machine Learning(慕尼黑机器学习中心) Technion - Israel Institute of Technology, Faculty of Mathematics(技术学院-以色列理工学院,数学学院)

专题命中 点云 :point cloud(abstract)

AI总结 本文提出自适应规范化解析,通过输入和网络共同决定规范形式,解决谱图神经网络的本征基模糊性和点云旋转对称性问题,实验表明其优于数据增强、标准规范化解析和等变架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03865 2026-04-17 cs.CG 50%

Towards an Optimal Bound for the Interleaving Distance on Mapper Graphs

迈向mapper图上交织距离的最优界限

Erin Wolf Chambers, Ishika Ghosh, Elizabeth Munch, Sarah Percival, Bei Wang

专题命中 点云 :point cloud(abstract)

AI总结 本文提出首个用于boundingmapper图上交织距离的框架,通过整数线性规划确定n-交织的存在性,并构造最小损失的assignment,用于分类任务。

Comments Reformulated the problem into both a binary problem and a loss computation problem

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 空间理解 3 篇

2604.11600 2026-04-17 cs.CV 57%

Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language

几何解析:一种统一形式语言用于平面和立体几何的图表解析

Peijie Wang, Ming-Liang Zhang, Jun Cao, Chao Deng, Dekang Ran, Hongda Sun, Pi Bu, Xuan Zhang, Yingyao Wang, Jun Song, Bo Zheng, Fei Yin, Cheng-Lin Liu

机构 * MAIS, Institute of Automation of Chinese Academy of Sciences(中国科学院自动化研究所MAIS) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Future Living Lab of Alibaba(阿里巴巴未来生活实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出一种统一形式语言,结合平面和立体几何,构建了GDP-29K数据集,通过监督微调与可验证奖励的强化学习提升几何解析性能,提升MLLMs的几何推理能力。

Comments Accepted to ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14641 2026-04-17 cs.AI 50%

Learning to Draw ASCII Improves Spatial Reasoning in Language Models

学习绘制ASCII图提高语言模型的空间推理能力

Shiyuan Huang, Li Liu, Jincheng He, Leilani H. Gilpin

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 空间理解 :spatial understanding(abstract)

AI总结 本文通过Text2Space数据集研究语言模型通过构造视觉布局提升空间推理能力,发现构造与理解训练结合可显著增强空间推理,并在外部基准上验证了其泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14637 2026-04-17 cs.HC 50%

Touching Space: Accessible Map Exploration Through Conversational Audio-Haptic Interaction

触碰空间:通过对话式音频-触觉交互实现可访问的地图探索

Li Liu, Jiaming Qu, Marc Jowell Bagaoisan, David T. Lee, Leilani H. Gilpin

专题命中 空间理解 :spatial understanding(abstract)

AI总结 本文提出Touching Space系统,通过触觉和音频反馈帮助视障人士构建环境认知地图,支持在普通硬件上实现空间探索。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. SLAM与定位 1 篇

2604.15052 2026-04-17 cs.RO 57%

CAVERS: Multimodal SLAM Data from a Natural Karstic Cave with Ground Truth Motion Capture

CAVERS:从自然喀斯特洞穴获取多模态SLAM数据并采用地面真实运动捕捉

Giacomo Franchini, David Rodríguez-Martínez, Alfonso Martínez-Petersen, C. J. Pérez-del-Pulgar, Marcello Chiaberge

机构 * Polytechnic of Turin Interdepartmental Centre for Service Robotics (PIC4SeR)(都灵理工大学服务机器人跨部门研究中心(PIC4SeR)) Systems Engineering and Automation Department, Universidad de Málaga(马德里大学系统工程与自动化系)

专题命中 SLAM与定位 :3D reconstruction(abstract);分类 cs.RO

AI总结 本文提出CAVERS数据集,通过地面真实运动捕捉系统提供高精度姿态和速度数据,验证了多模态传感器在复杂洞穴环境中的SLAM算法性能。

Comments 8 pages, 5 figures, preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏