arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 499 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 499 篇

2307.04760 2024-05-07 cs.CV cs.SD eess.AS 57%

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00962 2024-05-07 cs.CV cs.AI 57%

RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang, Xiaojuan Qi

专题命中 空间理解 :3D vision(abstract);分类 cs.CV

Comments To appear in CVPR2024 .project page: https://jihanyang.github.io/projects/RegionPLC

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11395 2024-04-23 cs.CV 57%

UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

Qingdong He, Jinlong Peng, Zhengkai Jiang, Kai Wu, Xiaozhong Ji, Jiangning Zhang, Yabiao Wang, Chengjie Wang, Mingang Chen, Yunsheng Wu

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Accepted by IJCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09502 2024-04-16 cs.CV 57%

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

Pin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments 10 pages, 4 figures, accepted by CVPR 2024

Journal ref IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06780 2024-04-11 cs.CV 57%

Urban Architect: Steerable 3D Urban Scene Generation with Layout Prior

Fan Lu, Kwan-Yee Lin, Yan Xu, Hongsheng Li, Guang Chen, Changjun Jiang

专题命中 空间理解 :3D generation(abstract);分类 cs.CV

Comments Project page: https://urbanarchitect.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06742 2024-04-02 cs.CV cs.AI cs.CL cs.LG 57%

Honeybee: Locality-enhanced Projector for Multimodal LLM

Junbum Cha, Wooyoung Kang, Jonghwan Mun, Byungseok Roh

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments CVPR 2024 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16431 2024-03-26 cs.CV cs.AI 57%

DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding

Xiaoxuan Yu, Hao Wang, Weiming Li, Qiang Wang, Soonyong Cho, Younghun Sung

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06547 2024-03-12 cs.CV 57%

AmodalSynthDrive: A Synthetic Amodal Perception Dataset for Autonomous Driving

Ahmed Rida Sekkat, Rohit Mohan, Oliver Sawade, Elmar Matthes, Abhinav Valada

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16308 2024-02-27 cs.RO 57%

DreamUp3D: Object-Centric Generative Models for Single-View 3D Scene Understanding and Real-to-Sim Transfer

Yizhe Wu, Haitz Sáez de Ocáriz Borde, Jack Collins, Oiwi Parker Jones, Ingmar Posner

专题命中 空间理解 :3D reconstruction(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09413 2024-01-18 cs.CV 57%

POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images

Antonin Vobecky, Oriane Siméoni, David Hurych, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Josef Sivic

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06758 2024-01-18 cs.RO cs.LG 57%

Exploring Contextual Representation and Multi-Modality for End-to-End Autonomous Driving

Shoaib Azam, Farzeen Munir, Ville Kyrki, Moongu Jeon, Witold Pedrycz

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11477 2023-11-21 cs.CV cs.CL 57%

What's left can't be right -- The remaining positional incompetence of contrastive vision-language models

Nils Hoehing, Ellen Rushe, Anthony Ventresque

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10076 2023-11-20 cs.CV 57%

A Simple Framework for 3D Occupancy Estimation in Autonomous Driving

Wanshui Gan, Ningkai Mo, Hongbin Xu, Naoto Yokoya

专题命中 空间理解 :3D reconstruction(abstract);分类 cs.CV

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13251 2023-10-31 cs.CV 57%

CGOF++: Controllable 3D Face Synthesis with Conditional Generative Occupancy Fields

Keqiang Sun, Shangzhe Wu, Ning Zhang, Zhaoyang Huang, Quan Wang, Hongsheng Li

专题命中 空间理解 :NeRF(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). This article is an extension of the NeurIPS'22 paper arXiv:2206.08361

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10015 2023-10-30 cs.CV cs.AI cs.CL 57%

Benchmarking Spatial Relationships in Text-to-Image Generation

Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vineet, Eric Horvitz, Ece Kamar, Chitta Baral, Yezhou Yang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments preprint; Code and Data at https://github.com/microsoft/VISOR and https://huggingface.co/datasets/tgokhale/sr2d_visor

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01091 2023-10-27 cs.CV cs.AI 57%

Changes to Captions: An Attentive Network for Remote Sensing Change Captioning

Shizhen Chang, Pedram Ghamisi

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09069 2023-10-16 cs.RO cs.AI 57%

ImageManip: Image-based Robotic Manipulation with Affordance-guided Next View Selection

Xiaoqi Li, Yanzi Wang, Yan Shen, Ponomarenko Iaroslav, Haoran Lu, Qianxu Wang, Boshi An, Jiaming Liu, Hao Dong

专题命中 空间理解 :point cloud(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10932 2023-09-21 cs.RO 57%

Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation

Tuan Van Vo, Minh Nhat Vu, Baoru Huang, Toan Nguyen, Ngan Le, Thieu Vo, Anh Nguyen

专题命中 空间理解 :point cloud(abstract);分类 cs.RO

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14727 2023-09-13 cs.CV 57%

You Only Need One Thing One Click: Self-Training for Weakly Supervised 3D Scene Understanding

Zhengzhe Liu, Xiaojuan Qi, Chi-Wing Fu

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Extension of One Thing One Click (CVPR'2021) arXiv:2104.02246

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12539 2023-08-09 cs.HC cs.GR 57%

VIRD: Immersive Match Video Analysis for High-Performance Badminton Coaching

Tica Lin, Alexandre Aouididi, Zhutian Chen, Johanna Beyer, Hanspeter Pfister, Jui-Hsien Wang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.GR

Comments To Appear in IEEE Transactions on Visualization and Computer Graphics (IEEE VIS), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10950 2023-06-22 cs.CV 57%

Factored Neural Representation for Scene Understanding

Yu-Shiang Wong, Niloy J. Mitra

专题命中 空间理解 :novel view synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15694 2023-05-26 cs.CV 57%

Learning Occupancy for Monocular 3D Object Detection

Liang Peng, Junkai Xu, Haoran Cheng, Zheng Yang, Xiaopei Wu, Wei Qian, Wenxiao Wang, Boxi Wu, Deng Cai

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11768 2023-05-26 cs.CV cs.CL 57%

Generating Visual Spatial Description via Holistic 3D Scene Understanding

Yu Zhao, Hao Fei, Wei Ji, Jianguo Wei, Meishan Zhang, Min Zhang, Tat-Seng Chua

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04926 2023-04-07 cs.CV 57%

CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP

Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li, Yuenan Hou, Yu Qiao, Wenping Wang

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08361 2022-12-13 cs.CV 57%

Controllable 3D Face Synthesis with Conditional Generative Occupancy Fields

Keqiang Sun, Shangzhe Wu, Zhaoyang Huang, Ning Zhang, Quan Wang, HongSheng Li

专题命中 空间理解 :NeRF(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08537 2022-10-18 cs.RO 57%

Learning 6-DoF Task-oriented Grasp Detection via Implicit Estimation and Visual Affordance

Wenkai Chen, Hongzhuo Liang, Zhaopeng Chen, Fuchun Sun, Jianwei Zhang

专题命中 空间理解 :point cloud(abstract);分类 cs.RO

Comments Accepted to the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07991 2022-10-17 cs.CV 57%

Novel 3D Scene Understanding Applications From Recurrence in a Single Image

Shimian Zhang, Skanda Bharadwaj, Keaton Kraiger, Yashasvi Asthana, Hong Zhang, Robert Collins, Yanxi Liu

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13156 2022-09-28 cs.CV cs.AI 57%

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi, Jian Zhang

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Submitted to ICRA2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03204 2022-07-08 cs.CV cs.AI cs.GT 57%

MCTS with Refinement for Proposals Selection Games in Scene Understanding

Sinisa Stekovic, Mahdi Rad, Alireza Moradi, Friedrich Fraundorfer, Vincent Lepetit

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Submitted to: TPAMI Special Section on the Best Papers of ICCV2021 GitHub Repository: https://github.com/vevenom/MonteScene. arXiv admin note: substantial text overlap with arXiv:2103.11161

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13410 2022-06-06 cs.CV 57%

KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D

Yiyi Liao, Jun Xie, Andreas Geiger

专题命中 空间理解 :novel view synthesis(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:1511.03240

详情

展开后加载摘要…

URL PDF HTML 收藏