Probabilistic Recovery of Multiple Subspaces in Point Clouds by Geometric lp Minimization
专题命中 点云 :point cloud(title)
Comments This paper was split into two different papers: 1. http://arxiv.org/abs/1012.4116 2. http://arxiv.org/abs/1104.3770
视觉与机器人
三维重建、NeRF、Gaussian Splatting、点云和空间智能。
专题命中 点云 :point cloud(title)
Comments This paper was split into two different papers: 1. http://arxiv.org/abs/1012.4116 2. http://arxiv.org/abs/1104.3770
专题命中 点云 :point cloud(title)
BridgeVLA++:一种面向三维操作的数据高效、可泛化且内存增强的视觉-语言-动作框架
机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室(NLPR)) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; FiveAges ; ByteDance Seed(字节跳动种子实验室)
专题命中 点云 :3D vision(abstract);point cloud(abstract);分类 cs.RO
AI总结 本研究提出内存增强的三维VLA框架BridgeVLA++,通过新增时空记忆架构,在保留原模型数据效率与泛化能力的同时,提升了记忆相关操作性能,且在多任务与真实平台上验证了其有效性。
Comments This work has been submitted to the IEEE TPAMI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
Wat3R:无需标注的水下3D几何学习
机构 * Huazhong University of Science and Technology(华中科技大学)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
AI总结 针对水下3D几何估计难题,提出跨域半监督学习框架Wat3R,基于师生架构,无需标注水下数据,利用未标注视频学习,设计跨视图一致性损失,构建Water3D数据集,实验证明其在水下多视图深度估计和点云重建上性能优于现有方法。
Comments Accepted to ECCV 2026. The dataset and code are available at https://github.com/LSXI7/Wat3R
用于学习鲁棒操纵策略的几何感知运动潜变量
机构 * Department of Computer Science, The University of Hong Kong(香港大学计算机科学系)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.RO
AI总结 研究如何学习机器人操纵的运动潜变量,核心方法是通过预测点云在操纵过程中的演变来学习离散运动潜码,贡献是仅用单视图RGB-D输入达先进性能,验证几何预测是关键,且潜码有有效运动抽象能力。
G2P:面向边界感知3D分割的高斯到点属性对齐
机构 * Kyungpook National University(庆北国立大学) ; Korea Electronics Technology Institute(韩国电子技术研究院) ; Adobe Research(Adobe研究院) ; Zhejiang University(浙江大学)
专题命中 点云 :Gaussian Splatting(abstract);point cloud(abstract);分类 cs.CV
AI总结 提出G2P方法,通过将3D高斯泼溅的属性(如不透明度和尺度)传递到点云,解决几何特征无法区分外观相似物体的问题,实现边界感知的3D分割。
Comments Accepted to ECCV 2026. Camera-ready version
MM-TRELLIS: 自动驾驶中基于点云引导的多模态3D车辆生成
机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室) ; Academy of Military Science(军事科学院) ; Bosch innovation software development (Wuxi) Co., Ltd.(博世创新软件开发(无锡)有限公司) ; College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院) ; Shopee Pte. Ltd.(Shopee私人有限公司)
专题命中 点云 :Gaussian Splatting(abstract);point cloud(abstract);分类 cs.CV
AI总结 提出MM-TRELLIS,融合多视图图像与LiDAR点云,通过点云引导和体素过滤策略,实现高质量3D车辆生成,在Waymo数据集上优于现有方法。
SpatialSV: 通过任务导向的视觉监督在多模态大语言模型中内化可解释的3D空间感知
机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能工程学院)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
AI总结 提出SpatialSV框架,通过任务导向的视觉监督将MLLM的2D特征提升为显式3D表示(深度图、相机姿态、点云),实现可解释的3D空间感知内化,无需外部工具,并在半监督设置中展现强泛化能力。
Comments Accepted by IJCAI 2026
S23DR 2026:基于对比去噪的DETR风格集合预测实现端到端3D线框预测
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
AI总结 提出WireframeDETR方法,直接对3D点云进行DETR风格集合预测,无需中间顶点检测,通过对比去噪训练、多尺度编码器和渐进辅助损失权重实现端到端3D线框预测,在S23DR 2026挑战赛上取得0.575 HSS。
Comments Technical report; S23DR 2026 Challenge submission
视觉与视觉-语言应用中的模态感知特征匹配:全面综述
机构 * School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(江西财经大学计算机与人工智能学院) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) ; School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) ; School of Computing and Communications, Lancaster University(兰卡斯特大学计算机与通讯学院) ; School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) ; Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(新加坡资讯研究院,科技研究局(A*STAR)) ; Department of Automation, Tsinghua University(清华大学自动化系)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
AI总结 综述基于模态的特征匹配,涵盖传统手工方法和现代深度学习方法,重点讨论跨RGB、深度、3D点云、LiDAR、医学图像及视觉-语言模态的进展,突出模态感知技术。
Comments CSUR
双路径几何感知多模态大语言模型用于空间智能
机构 * University of Science and Technology of China(中国科学技术大学) ; Li Auto Inc.(利汽车公司) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 提出GAMSI,一种仅以RGB图像为输入、通过双路径查询和专家引导视觉对齐实现3D结构与度量尺度联合感知的多模态大语言模型,在七个空间智能基准上达到最优性能。
面向自动驾驶的空间感知视觉语言模型
机构 * Motional ; University of Amsterdam(阿姆斯特丹大学)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 提出LVLDrive框架,通过融合LiDAR点云与视觉语言模型,利用渐进融合Q-Former和空间感知问答数据集,解决3D度量空间推理瓶颈,提升自动驾驶场景理解与决策可靠性。
Comments Accepted to CVPR AutoPilot Workshop 2026
Sensor2Sensor: 自动驾驶的跨本体传感器转换
机构 * Waymo ; Johns Hopkins University(约翰霍普金斯大学) ; Google DeepMind(谷歌DeepMind) ; University of Washington(华盛顿大学)
专题命中 点云 :Gaussian Splatting(abstract);point cloud(abstract);分类 cs.CV
AI总结 提出Sensor2Sensor生成模型,将单目行车记录仪视频转换为多模态传感器数据(多视角相机图像和LiDAR点云),通过4D高斯泼溅重建和扩散架构解决无配对数据问题,为自动驾驶开发解锁外部数据源。
Comments Accepted by CVPR 2026
扩展以占据为中心的驾驶场景生成:数据集与方法
机构 * Shanghai Jiao Tong University(上海交通大学) ; Eastern Institute of Technology(东部技术研究院) ; School of Electronic Information and Electrical Engineering(电子信息与电气工程学院) ; Li Auto(力汽车) ; National University of Singapore(新加坡国立大学) ; Tsinghua University(清华大学) ; Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生实验室) ; Ningbo Institute of Digital Twin(宁波数字孪生研究院)
专题命中 点云 :Gaussian Splatting(abstract);point cloud(abstract);分类 cs.CV
AI总结 针对占据数据稀缺问题,构建最大语义占据数据集Nuplan-Occ,并提出统一框架联合生成高质量语义占据、多视角视频和LiDAR点云,采用时空解耦架构及高斯泼溅稀疏点图渲染和传感器感知嵌入策略,实现高保真生成。
Comments IEEE TPAMI
通过运动诱导采样用消费级LiDAR成像隐藏物体
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Dartmouth College(达特茅斯学院)
专题命中 点云 :3D reconstruction(abstract,abstract_cn);分类 cs.CV
AI总结 本文提出了一种多帧融合策略,利用运动诱导孔径采样模型,在消费级LiDAR上实现了非线视成像,实现了隐藏物体的3D重建、多物体跟踪和相机定位,并展示了消费级硬件无需额外设置即可实现非线视成像的潜力。
低成本立体视觉用于自主无人机修剪中薄辐射松枝的稳健三维定位
机构 * Centre of Data Science and Artificial Intelligence & School of Engineering and Computer Science(数据科学与人工智能中心及工程与计算机科学学院)
专题命中 点云 :NeRF(abstract,abstract_cn);分类 cs.CV
AI总结 本文研究利用低成本立体相机实现自主无人机修剪薄枝的三维定位,通过分支分割与深度估计方法,解决森林场景中纹理稀疏、结构细长和噪声干扰等问题。
如何使LLM具备3D能力?LLM中空间推理的综述
机构 * Tsinghua University(清华大学) ; The Hong Kong University of Science and Technology (Guang Zhou)(香港科学与技术大学(广州))
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 本文综述了将LLM与3D空间理解结合的方法,提出分类体系,涵盖图像、点云和混合模态方法,并讨论了当前限制与未来研究方向。
Comments 9 pages, 5 figures
Journal ref IJCAI 2025
GGPT:基于几何的点变换
机构 * ETH Zurich(苏黎世联邦理工学院) ; Delft University of Technology(代尔夫特理工大学)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
AI总结 GGPT通过引入几何引导的点变换,结合密集前馈预测与显式几何约束,实现了更准确且一致的3D重建,提升了细粒度结构恢复和无纹理区域填补能力。
Comments CVPR 2026, Project website: https://chenyutongthu.github.io/research/ggpt
迈向自动驾驶的统一多模态表示学习
机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA(德克萨斯大学机械工程系,德克萨斯农工大学,学院站,德克萨斯,77843,美国) ; The Department of Engineering Technology and Industrial Distribution Texas A&M University, College Station, TX 77843, USA(工程技术与工业分布系,德克萨斯农工大学,学院站,德克萨斯,77843,美国)
专题命中 点云 :3D vision(abstract);point cloud(abstract);分类 cs.CV
AI总结 本文提出CTP框架,通过统一多模态张量对齐提升自动驾驶性能。
OnlineSI: 通过大规模语言模型实现在线3D理解和定位
机构 * Tsinghua University(清华大学) ; Nanyang Technological University(南洋理工大学) ; Shanghai AI Lab(上海人工智能实验室)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 OnlineSI通过整合3D点云与语义信息,提升大规模语言模型在动态环境中的空间理解和物体识别能力。
Comments Project Page: https://onlinesi.github.io/
基于3D几何先验的双臂操作动作预测
机构 * College of Computer Science, Sichuan University, China(四川大学计算机学院) ; University of Electronic Science and Technology of China, China(电子科技大学) ; Dexmal ; The Chinese University of Hong Kong, HK SAR, China(香港中文大学(HK SAR))
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 本文提出基于3D几何先验的双臂操作预测框架,通过融合几何特征与2D语义信息,实现高效动作预测与空间理解。
Comments Accepted by CVPR 2026
自动驾驶中的3D物体检测:综述
机构 * Key Lab of Data Engineering and Knowledge Engineering, Renmin University of China, Beijing 100872, China(数据工程与知识工程重点实验室,中国人民大学,北京100872,中国) ; School of Mathematics, Renmin University of China, Beijing 100872, China(数学学院,中国人民大学,北京100872,中国)
专题命中 点云 :3D vision(abstract);point cloud(abstract);分类 cs.CV
AI总结 本文综述了自动驾驶中3D物体检测的研究进展,涵盖了传感器、数据集、性能指标及最新方法,并分析了其优缺点及未来方向。
Comments The manuscript is accepted by Pattern Recognition on 14 May 2022
Zoo3D:面向场景级别的零样本3D物体检测
机构 * Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学) ; Higher School of Economics(高等经济学院) ; M:3L Lab, Institute of Mechanics, Armenia(M:3L实验室,力学研究所,亚美尼亚)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
AI总结 Zoo3D提出无需训练的3D物体检测框架,通过图聚类和开放式词汇模块实现零样本和自监督模式,取得开放词汇3D检测最佳效果。
机构 * University of California, Los Angeles(加州大学洛杉矶分校) ; University of pennsylvania(宾夕法尼亚大学)
专题命中 点云 :3D vision(abstract);point cloud(abstract);分类 cs.RO
Comments NeurIPS 2025 Workshop on Embodied World Models for Decision Making. URL: https://avi-3drobot.github.io/
机构 * Westlake University(西湖大学) ; Zhejiang University(浙江大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * Xiaomi EV(小米电动车)
专题命中 点云 :point cloud(abstract);novel view synthesis(abstract);分类 cs.CV
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 点云 :point cloud(abstract);spatial understanding(abstract);分类 cs.CV
Comments Code and data are available at https://zoezheng126.github.io/STLLM-website/
机构 * Department of Computer Science and Engineering, Lehigh University(计算机科学与工程系,莱维大学) ; Department of Civil and Environmental Engineering, Lehigh University(土木与环境工程系,莱维大学)
专题命中 点云 :3D vision(abstract);point cloud(abstract);分类 cs.CV
机构 * University of Melbourne(墨尔本大学) ; KU Leuven(鲁文大学)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
Comments 5 pages, 4 figures
Journal ref Published on 2025 IEEE International Geoscience and Remote Sensing Symposium
机构 * Keio University(Keio大学) ; AISIN CORPORATION(AISIN公司)
专题命中 点云 :3D reconstruction(abstract);point cloud(abstract);分类 cs.CV
Comments Accepted to MMSports