arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 7719 信号源:cs.CV, cs.GR, cs.RO

1. 点云 7719 篇

2601.13364 2026-01-21 cs.CV 57%

Real-Time 4D Radar Perception for Robust Human Detection in Harsh Enclosed Environments

实时4D雷达感知用于恶劣封闭环境中的可靠人类检测

Zhenan Liu, Yaodong Cui, Amir Khajepour, George Shaker

机构 * University of Waterloo(滑铁卢大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种实时4D毫米波雷达感知方法,通过噪声过滤和基于规则的分类流程,在尘埃环境中实现可靠的人类检测。

Journal ref 2025 IEEE International Symposium on Antennas and Propagation and North American Radio Science Meeting (AP-S/CNC-USNC-URSI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06467 2026-01-21 cs.CV 57%

Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration

DINOv3 是否设定了医学视觉的新标准?对2D和3D分类、分割与配准的基准测试

Che Liu, Yinda Chen, Haoyuan Shi, Jinpeng Lu, Bailiang Jian, Jiazhen Pan, Linghan Cai, Jiayi Wang, Jieming Yu, Ziqi Gao, Xiaoran Zhang, Long Bai, Yundi Zhang, Jun Li, Cosmin I. Bercea, Cheng Ouyang, Chen Chen, Zhiwei Xiong, Benedikt Wiestler, Christian Wachinger, James S. Duncan, Daniel Rueckert, Wenjia Bai, Rossella Arcucci

机构 * Imperial College London(伦敦帝国理工学院) University of Science and Technology of China(中国科学技术大学) Dresden University of Technology(德累斯顿技术大学) University of Erlangen-Nuremberg(埃尔兰根-纽伦堡大学) University of Oxford(牛津大学) University of Sheffield(谢菲尔德大学) Technical University of Munich (TUM)(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学) Yale University(耶鲁大学)

专题命中 点云 :3D reconstruction(abstract);分类 cs.CV

AI总结 DINOv3在医学视觉任务中表现出色,但其在深度领域专门化任务中存在性能退化问题。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11269 2026-01-19 cs.CV cs.AI 57%

X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning

X-Distill:跨架构视觉蒸馏用于视觉-运动学习

Maanping Shao, Feihong Zhang, Gu Zhang, Baiye Cheng, Zhengrong Xue, Huazhe Xu

机构 * Tsinghua University(清华大学) Institute for Interdisciplinary Information Sciences(交叉信息研究院) Shanghai Qi Zhi Institute(上海启智研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 X-Distill通过跨架构知识蒸馏结合视觉转换器和紧凑型CNN,实现了在数据有限的机器人操作任务中优于其他方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11354 2026-01-15 cs.CV 57%

A Multi-Mode Structured Light 3D Imaging System with Multi-Source Information Fusion for Underwater Pipeline Detection

一种基于多源信息融合的多模式结构光3D成像系统用于水下管道检测

Qinghan Hu, Haijiang Zhu, Na Sun, Lei Chen, Zhengqiang Fan, Zhiqing Li

机构 * College of Information Science and Technology, Beijing University of Chemical Technology, Beijing 100029, China(信息科学与技术学院,北京化工大学,中国北京100029) Department of Mechanical Engineering, Tsinghua University, Beijing 100084, China(机械工程系,清华大学,中国北京100084) Guoneng Zhishen Control Technology Co., Ltd., Beijing 102211, China(国能智深控制技术有限公司,中国北京102211) College of Intelligent Science and Engineering, Beijing University of Agriculture, Beijing 102206, China(智能科学与工程学院,北京农业大学,中国北京102206)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种基于多源信息融合的多模式结构光3D成像系统,用于水下管道检测,通过改进的校正方法和融合策略提升检测精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10046 2026-01-15 cs.CV 57%

A Preprocessing and Postprocessing Voxel-based Method for LiDAR Semantic Segmentation Improvement in Long Distance

一种用于增强长距离LiDAR语义分割的预处理和后处理体素方法

Andrea Matteazzi, Pascal Colling, Michael Arnold, Dietmar Tutsch

机构 * University of Wuppertal(乌尔姆大学) Aptiv Services Deutschland GmbH(APTIV德国服务公司)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种用于增强长距离LiDAR语义分割的预处理和后处理方法,通过多阶段处理和先进模型结合,显著提升了中远距离的mIoU性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07119 2026-01-13 cs.DC cs.CV 57%

SC-MII: Infrastructure LiDAR-based 3D Object Detection on Edge Devices for Split Computing with Multiple Intermediate Outputs Integration

SC-MII:基于基础设施激光雷达的边缘设备上用于分割计算的多中间输出集成的3D目标检测

Taisuke Noguchi, Takayuki Nishio, Takuya Azumi

机构 * Graduate School of Science and Engineering(科学与工程研究生学校) Saitama University(上野大学) School of Engineering(工程学院) Institute of Science Tokyo(东京科学研究所)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 SC-MII通过多激光雷达和边缘计算实现高效3D目标检测,降低延迟和能耗,提升隐私保护。

Comments 6 pages. This version includes minor lstlisting configuration adjustments for successful compilation. No changes to content or layout. Originally published at IEEE CCNC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03519 2026-01-13 cs.RO 57%

A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving

一种具有视觉提示的视觉-语言-动作模型用于越野自动驾驶

Liangdong Zhang, Yiming Nie, Haoyang Li, Fanjie Kong, Baobao Zhang, Shunxin Huang, Kai Fu, Chen Min, Liang Xiao

专题命中 点云 :spatial understanding(abstract);分类 cs.RO

AI总结 本文提出OFF-EMMA模型,通过视觉提示和COT-SC策略提升越野自动驾驶轨迹规划的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06496 2026-01-13 cs.CV 57%

3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence

3D CoCa v2:基于测试时间搜索的对比学习用于通用空间智能

Hao Tang, Ting Huang, Zeyu Zhang

机构 * School of Computer Science, Peking University(北京大学计算机学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 3D CoCa v2通过测试时间搜索提升3D场景描述的泛化能力,实现对比学习与描述生成的统一,提高鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06465 2026-01-13 eess.IV cs.CV cs.MM 57%

R$^3$D: Regional-guided Residual Radar Diffusion

R$^3$D: 基于区域引导的残差雷达扩散

Hao Li, Xinqi Liu, Yaoqing Jin

机构 * University of Arizona(亚利桑那大学) University of Hong Kong(香港大学) University of Stuttgart(斯图加特大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 R3D通过区域引导残差雷达扩散框架,整合残差建模与sigma自适应引导,提升雷达点云质量,优于现有方法。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06464 2026-01-13 cs.CV 57%

On the Adversarial Robustness of 3D Large Vision-Language Models

关于3D大视觉-语言模型的对抗鲁棒性

Chao Liu, Ngai-Man Cheung

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 点云 :3D vision(abstract);分类 cs.CV

AI总结 本文研究了3D大视觉-语言模型的对抗鲁棒性,提出两种攻击策略评估其鲁棒性,发现其在无目标攻击下较脆弱,但在有目标攻击中更稳健。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19290 2026-01-13 cs.CV 57%

TRASE: Tracking-free 4D Segmentation and Editing

TRASE:无需跟踪的4D分割与编辑

Yun-Jin Li, Mariia Gladkova, Yan Xia, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 点云 :3D reconstruction(abstract);分类 cs.CV

AI总结 TRASE通过弱监督学习实现无需跟踪的4D分割,利用对比学习和聚类技术实现动态场景的高效分割与交互编辑。

Comments Accepted to 3DV 2026. Project page https://yunjinli.github.io/project-sadg

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19500 2026-01-07 cs.RO cs.LG 57%

RobotDiffuse: Diffusion-Based Motion Planning for Redundant Manipulators with the ROP Obstacle Avoidance Dataset

RobotDiffuse: 基于扩散模型的冗余机械臂运动规划与ROP障碍避让数据集

Xudong Mou, Xiaohan Zhang, Tiejun Wang, Tianyu Wo, Cangbai Xu, Ningbo Gu, Rui Wang, Xudong Liu

机构 * School of Computer Science and Engineering, Beihang University, Beijing, China(北京航空航天大学计算机科学与工程学院) School of Software, Beihang University, Beijing, China(北京航空航天大学软件学院) Hangzhou Innovation Institute, Beihang University, Hangzhou, China(北京航空航天大学杭州创新研究院)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 RobotDiffuse通过扩散模型实现冗余机械臂运动规划,结合物理约束与点云编码器,并发布ROP数据集提升障碍避让性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02392 2026-01-07 cs.CV cs.AI 57%

Self-Supervised Masked Autoencoders with Dense-Unet for Coronary Calcium Removal in limited CT Data

具有密集-UNet的自监督掩码自动编码器用于有限CT数据中的冠状动脉钙化去除

Mo Chen

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出Dense-MAE,一种基于自监督学习的框架,用于在有限CT数据中有效去除冠状动脉钙化伪影,提升血管狭窄诊断准确性。

Comments 6 pages, in Chinese language, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02029 2026-01-06 cs.CV 57%

Leveraging 2D-VLM for Label-Free 3D Segmentation in Large-Scale Outdoor Scene Understanding

利用2D-VLM实现无标注3D分割以支持大规模户外场景理解

Toshihiko Nishimura, Hirofumi Abe, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida

机构 * NTT Corporation(日本电报电话公司)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出基于2D-VLM的无标注3D分割方法,通过虚拟相机投影和自然语言提示实现大规模户外场景的语义分割,支持开放词汇识别。

Comments 19

Journal ref 19th International Conference on Machine Vision Applications (MVA2025), IEICE Transactions on Information and Systems letter

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01044 2026-01-06 cs.CV cs.LG 57%

Evaluating transfer learning strategies for improving dairy cattle body weight prediction in small farms using depth-image and point-cloud data

评估用于改进小型农场奶牛体重预测的迁移学习策略:使用深度图像和点云数据

Jin Wang, Angelo De Castro, Yuxi Zhang, Lucas Basolli Borsatto, Yuechen Guo, Victoria Bastos Primo, Ana Beatriz Montevecchio Bernardino, Gota Morota, Ricardo C Chebel, Haipeng Yu

机构 * Department of Animal Sciences, University of Florida(动物科学系,佛罗里达大学) Department of Large Animal Clinical Sciences, University of Florida(大动物临床科学系,佛罗里达大学) Laboratory of Biometry and Bioinformatics, Department of Agricultural and Environmental Biology, Graduate School of Agricultural and Life Sciences, The University of Tokyo(生物度量与生物信息学实验室,农业与环境生物学系,东京大学研究生院) Department of Agricultural and Environmental Biology, Graduate School of Agricultural and Life Sciences, The University of Tokyo(农业与环境生物学系,东京大学研究生院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本研究评估了迁移学习在小型农场奶牛体重预测中的有效性,比较了深度图像和点云数据的预测性能,发现迁移学习在不同农场条件下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21890 2026-01-05 cs.CV 57%

CrownGen: Patient-customized Crown Generation via Point Diffusion Model

CrownGen: 通过点扩散模型实现患者定制的牙冠生成

Juyoung Bae, Moo Hyun Son, Jiale Peng, Wanting Qu, Wener Chen, Zelin Qiu, Kaixin Li, Xiaojuan Chen, Yifan Lin, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Hong Kong University of Science and Technology(香港科学与技术大学) Division of Pediatric Dentistry and Orthodontics, Faculty of Dentistry(牙科学院儿童牙科与正畸科) University of Hong Kong(香港大学) Department of Chemical and Biological Engineering(化学与生物工程系) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港联合创新研究院) State Key Laboratory of Nervous System Disorders(神经系统疾病国家重点实验室) Delun Dental Hospital(德尔润牙科医院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 CrownGen通过点扩散模型实现患者定制的牙冠自动化生成,提升了牙科修复的效率和质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24605 2026-01-01 cs.CV 57%

MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding

MoniRefer: 一种基于道路基础设施的现实世界大规模多模态数据集用于3D视觉定位

Panquan Yang, Junfei Huang, Zongzhangbao Yin, Yingsong Hu, Anni Xu, Xinyi Luo, Xueqi Sun, Hai Wu, Sheng Ao, Zhaoxing Zhu, Chenglu Wen, Cheng Wang

机构 * Xiamen University(厦门大学) Pengcheng Laboratory(鹏城实验室)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出MoniRefer数据集和Moni3DVG方法,用于解决户外监控场景下的3D视觉定位问题。

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24603 2026-01-01 cs.CV 57%

Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers

协同低秩适应用于预训练视觉变换器

Zheng Liu, Jinchao Zhu, Gao Huang

机构 * School of Automation and Electrical Engineering, University of Science and Technology Beijing(北京科技大学自动化与电气工程学院) Beijing Engineering Research Center of Industrial Spectrum Imaging(工业光谱成像工程研究中心) College of Software, Nankai University(南开大学软件学院) Department of Automation, BNRist, Tsinghua University(清华大学自动化系)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出协同低秩适应方法,通过基础空间共享和样本无关多样性增强,提升预训练视觉变换器在点云分析中的性能与参数效率。

Comments 13 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24593 2026-01-01 cs.CV cs.LG 57%

3D Semantic Segmentation for Post-Disaster Assessment

灾害评估中的3D语义分割

Nhut Le, Maryam Rahnemoonfar

机构 * Lehigh University(莱文斯顿大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出利用无人机拍摄的灾后影像构建专用3D数据集,并评估现有模型在灾后场景中的局限性,强调了改进3D分割技术与开发专用基准数据集的必要性。

Comments Accepted by the 2025 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23972 2026-01-01 cs.RO 57%

SHIELD: Spherical-Projection Hybrid-Frontier Integration for Efficient LiDAR-based Drone Exploration

SHIELD:球面投影混合前沿集成用于高效的基于LiDAR的无人机探索

Liangtao Feng, Zhenchang Liu, Feng Zhang, Xuefeng Ren

机构 * Zhuoyi Intelligent Tech Co, Ltd.(卓奕智能科技有限公司)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 SHIELD通过球面投影和混合前沿方法提升基于LiDAR的无人机探索效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22304 2025-12-30 cs.CV 57%

PortionNet: Distilling 3D Geometric Knowledge for Food Nutrition Estimation

PortionNet:基于3D几何知识的食品营养估计

Darrin Bright, Rakshith Raj, Kanchan Keisham

机构 * Vellore Institute of Technology(维捷里尼学院)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 PortionNet通过跨模态知识蒸馏框架,在无需专用硬件的情况下,利用点云学习几何特征,实现单图像食品营养估计的高精度

Comments Accepted at the 11th Annual Conference on Vision and Intelligent Systems (CVIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20217 2025-12-24 cs.CV 57%

LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation

LiteFusion: 从基于视觉到多模态的3D目标检测器的最小适应

Xiangxuan Ren, Zhongdao Wang, Pin Tang, Guoqing Wang, Jilai Zheng, Chao Ma

机构 * China Ministry of Education (MOE) Key Laboratory of Artificial Intelligence(中国教育部人工智能重点实验室) Artificial Intelligence Institute, Shanghai Jiao Tong University(上海交通大学人工智能学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 LiteFusion通过将LiDAR数据作为补充几何信息源,无需专用LiDAR编码器,显著提升多模态3D目标检测性能。

Comments 13 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17085 2025-12-24 cs.RO cs.LG 57%

Deformable Cluster Manipulation via Whole-Arm Policy Learning

通过全身政策学习实现可变形簇 manipulation

Jayadeep Jacob, Wenzheng Zhang, Houston Warren, Paulo Borges, Tirthankar Bandyopadhyay, Fabio Ramos

机构 * School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) Data61, CSIRO(CSIRO数据61研究所) Orica(奥里卡公司) NVIDIA Corporation(NVIDIA公司)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出了一种基于全身政策学习的方法,通过整合3D点云和本体感觉触觉信息,实现对可变形簇的高效操纵,并在输电线路清除任务中展示了零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05634 2025-12-24 cs.CV cs.LG eess.IV 57%

Fusarium head blight detection, spikelet estimation, and severity assessment in wheat using 3D convolutional neural networks

利用3D卷积神经网络检测小麦 Fusarium head blight、估计穗数及评估病情严重程度

Oumaima Hamila, Christopher J. Henry, Oscar I. Molina, Christopher P. Bidinosti, Maria Antonia Henriquez

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本研究利用3D卷积神经网络对小麦FHB进行检测、穗数估计及严重程度评估,实现100%检测准确率和高精度估计

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19088 2025-12-23 cs.CV 57%

Retrieving Objects from 3D Scenes with Box-Guided Open-Vocabulary Instance Segmentation

通过框引导的开放词汇实例分割从3D场景中检索物体

Khanh Nguyen, Dasith de Silva Edirimuni, Ghulam Mubashar Hassan, Ajmal Mian

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出了一种基于2D开放词汇检测器引导的3D实例分割方法,用于从RGB图像中快速准确检索稀有物体实例。

Comments Accepted to AAAI 2026 Workshop on New Frontiers in Information Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17764 2025-12-22 cs.RO 57%

UniStateDLO: Unified Generative State Estimation and Tracking of Deformable Linear Objects Under Occlusion for Constrained Manipulation

UniStateDLO:统一的生成式状态估计与变形线性物体在遮挡下的跟踪用于受限操作

Kangchen Lv, Mingrui Yu, Shihefeng Wang, Xiangyang Ji, Xiang Li

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 UniStateDLO通过深度学习方法实现可变形线性物体在遮挡下的鲁棒状态估计与跟踪,利用扩散模型提升复杂映射的捕捉能力,展现高效的数据利用和仿真到现实的泛化能力。

Comments The first two authors contributed equally. Project page: https://unistatedlo.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17568 2025-12-22 cs.RO 57%

Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation

具有一致3D观察和动作空间的运动学感知扩散策略用于全身机械臂操作

Kangchen Lv, Mingrui Yu, Yongyi Jia, Chenyu Zhang, Xiang Li

专题命中 点云 :point cloud(abstract);分类 cs.RO

AI总结 本文提出一种基于一致3D空间的运动学感知扩散策略,用于提高全身机械臂操作的样本效率和空间泛化能力。

Comments The first two authors contributed equally. Project Website: https://kinematics-aware-diffusion-policy.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05144 2025-12-22 cs.CV 57%

SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing

SGS-3D: 通过可靠的语义掩码分割与生长实现高保真的3D实例分割

Chaolei Wang, Yang Luo, Jing Du, Siyu Chen, Yiping Chen, Ting Han

专题命中 点云 :3D vision(abstract);分类 cs.CV

AI总结 SGS-3D通过可靠的语义掩码分割与生长方法,提升3D实例分割的精度和鲁棒性,适用于多样化的室内和室外环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15716 2025-12-18 cs.CV cs.AI 57%

Spatia: Video Generation with Updatable Spatial Memory

基于可更新空间记忆的视频生成

Jinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang, Chang Xu, Yan Lu

机构 * The University of Sydney(悉尼大学) Microsoft Research(微软研究院) HKUST(香港科技大学) University of Waterloo(滑铁卢大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 Spatia通过可更新的空间记忆机制,实现视频生成中的长期空间与时间一致性,支持动态实体生成和3D交互编辑。

Comments Project page: https://zhaojingjing713.github.io/Spatia/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17467 2025-12-17 cs.CV 57%

Semantic-Free Procedural 3D Shapes Are Surprisingly Good Teachers

无语义的程序化3D形状出人意料地成为好的教师

Xuweiyi Chen, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学)

专题命中 点云 :point cloud(abstract);分类 cs.CV

AI总结 本文提出通过程序化生成的3D形状学习3D表示,发现其在多个下游任务中表现优异,表明3D自监督学习不依赖语义信息。

Comments 3DV | SynData4CV @ CVPR2025 | Project Page: https://point-mae-zero.cs.virginia.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏