arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 499 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 499 篇

2506.12214 2025-06-17 cs.CV 57%

CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images

Ilya Ilyankou, Natchapon Jongwiriyanurak, Tao Cheng, James Haworth

机构 * UCL SpaceTimeLab(伦敦大学空间时间实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09565 2025-06-16 cs.CV 57%

SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields

Qijing Li, Jingxiang Sun, Liang An, Zhaoqi Su, Hongwen Zhang, Yebin Liu

机构 * Beijing Normal University(北京师范大学) Tsinghua University(清华大学)

专题命中 空间理解 :3D reconstruction(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02690 2025-06-04 cs.CV math.GT 57%

Towards Geometry Problem Solving in the Large Model Era: A Survey

Yurui Zhao, Xiang Wang, Jiahong Liu, Irwin King, Zhitao Huang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments 8pages, 4 figures, conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21079 2025-05-28 cs.CV 57%

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Yue Zhang, Yingzhao Jian, Hehe Fan, Yi Yang, Roger Zimmermann

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20640 2025-05-28 cs.CV 57%

IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios

Yifan Li, Yuhang Chen, Anh Dao, Lichi Li, Zhongyi Cai, Zhen Tan, Tianlong Chen, Yu Kong

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments v1.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12589 2025-05-20 cs.CV 57%

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Bo Liu, Pengfei Qiao, Minhan Ma, Xuange Zhang, Yinan Tang, Peng Xu, Kun Liu, Tongtong Yuan

机构 * Department of Computer Science, Beijing University of Technology(北京理工大学计算机科学系) Inspur Electronic Information Industry Co., Ltd(Inspur电子信息产业有限公司) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) JD Explore Academy(京东探索研究院)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments The dataset and code are publicly available at: https://huggingface.co/datasets/fei213/SurveillanceVQA-589K

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15830 2025-05-20 cs.RO cs.AI 57%

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, Xuelong Li

机构 * Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) ShanghaiTech University(上海科技大学) Northwestern Polytechnical University(西北工业大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

Journal ref Robotics: Science and Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13580 2025-05-19 cs.CV 57%

Leveraging Automatic CAD Annotations for Supervised Learning in 3D Scene Understanding

Yuchen Rao, Stefan Ainetter, Sinisa Stekovic, Vincent Lepetit, Friedrich Fraundorfer

机构 * Inst. of Visual Computing, Graz Univ. of Technology, Austria(视觉计算研究所,格拉茨技术大学,奥地利) LIGM, École des Ponts et Chaussees, IP Paris, CNRS, France(LIGM,巴黎理工学院,IP巴黎,CNRS,法国)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Project page: https://stefan-ainetter.github.io/SCANnotatepp; CVPR'25 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21696 2025-05-15 cs.CL cs.CV 57%

Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Wenqi Zhang, Mengna Wang, Gangao Liu, Xu Huixin, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Weiming Lu, Peng Li, Yueting Zhuang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) Alibaba Group(阿里巴巴集团) DAMO Academy, Alibaba Group(阿里巴巴集团大模型学院) Nanjing Institute of Software Technology(南京软件技术研究所) Nanjing University of Posts and Telecommunications(南京邮电大学) Hohai University(河海大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments Code: https://github.com/zwq2018/embodied_reasoner Dataset: https://huggingface.co/datasets/zwq2018/embodied_reasoner

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06136 2025-05-12 cs.RO cs.AI 57%

Efficient Sensorimotor Learning for Open-world Robot Manipulation

Yifeng Zhu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

Comments Ph.D. Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19500 2025-04-29 cs.CV cs.CL 57%

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

Yan Wang, Baoxiong Jia, Ziyu Zhu, Siyuan Huang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19266 2025-04-29 cs.CV 57%

OpenFusion++: An Open-vocabulary Real-time Scene Understanding System

Xiaofeng Jin, Matteo Frosi, Matteo Matteucci

机构 * Politecnico di Milano(米兰理工学院)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18125 2025-04-29 cs.CV 57%

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 空间理解 :3D vision(abstract);分类 cs.CV

Comments Project page: https://zcmax.github.io/projects/LLaVA-3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11208 2025-04-24 cs.CV cs.LG 57%

PooDLe: Pooled and dense self-supervised learning from naturalistic videos

Alex N. Wang, Christopher Hoang, Yuwen Xiong, Yann LeCun, Mengye Ren

机构 * New York University(纽约大学) Meta

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments Project page: https://agenticlearning.ai/poodle/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06141 2025-04-22 cs.RO cs.HC 57%

Mixed Reality Outperforms Virtual Reality for Remote Error Resolution in Pick-and-Place Tasks

Advay Kumar, Stephanie Simangunsong, Pamela Carreno-Medrano, Akansel Cosgun

机构 * Faculty of Engineering(工程学院) Monash University(墨尔本大学) Faculty of Science Engineering and Built Environment(科学工程与建筑环境学院) Deakin University(迪金大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09583 2025-04-18 cs.RO 57%

ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Runyu Ma, Jelle Luijkx, Zlatan Ajanovic, Jens Kober

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

Comments 6 pages, 6 figures, IEEE International Conference on Robotics and Automation (ICRA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10106 2025-04-15 cs.CV cs.AI 57%

SoccerNet-v3D: Leveraging Sports Broadcast Replays for 3D Scene Understanding

Marc Gutiérrez-Pérez, Antonio Agudo

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02971 2025-04-09 cs.CV cs.CL 57%

QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding

Binh M. Le, Shaoyuan Xu, Jinmiao Fu, Zhishen Huang, Moyan Li, Yanhui Guo, Hongdong Li, Sameera Ramasinghe, Bryan Wang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments 8 pages, accepted by CVPR 2025 MULA

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02328 2025-04-04 cs.CV 57%

Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

Congpei Qiu, Yanhao Wu, Wei Ke, Xiuxiu Bai, Tong Zhang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16137 2025-04-02 cs.RO cs.AI 57%

A Graph-to-Text Approach to Knowledge-Grounded Response Generation in Human-Robot Interaction

Nicholas Thomas Walker, Stefan Ultes, Pierre Lison

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

Comments Submitted to Dialogue & Discourse 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22168 2025-03-31 cs.CV 57%

Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis

Woojung Han, Yeonkyung Lee, Chanyoung Kim, Kwanghyun Park, Seong Jae Hwang

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01550 2025-03-24 cs.CV cs.AI 57%

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

Chunlin Yu, Hanqing Wang, Ye Shi, Haoyang Luo, Sibei Yang, Jingyi Yu, Jingya Wang

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03142 2025-03-21 cs.RO 57%

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

Shijie Wu, Yihang Zhu, Yunao Huang, Kaizhen Zhu, Jiayuan Gu, Jingyi Yu, Ye Shi, Jingya Wang

专题命中 空间理解 :point cloud(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13924 2025-03-21 cs.CV cs.AI 57%

ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding

Guangda Ji, Silvan Weder, Francis Engelmann, Marc Pollefeys, Hermann Blum

专题命中 空间理解 :3D vision(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10982 2025-03-17 cs.CV 57%

Enhanced Multi-View Pedestrian Detection Using Probabilistic Occupancy Volume

Reef Alturki, Adrian Hilton, Jean-Yves Guillemaut

专题命中 空间理解 :3D reconstruction(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09941 2025-03-14 cs.CV cs.AI 57%

TGP: Two-modal occupancy prediction with 3D Gaussian and sparse points for 3D Environment Awareness

Mu Chen, Wenyu Chen, Mingchuan Yang, Yuan Zhang, Tao Han, Xinchi Li, Yunlong Li, Huaici Zhao

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07125 2025-03-11 cs.CV 57%

Learning A Zero-shot Occupancy Network from Vision Foundation Models via Self-supervised Adaptation

Sihao Lin, Daqi Liu, Ruochong Fu, Dongrui Liu, Andy Song, Hongwei Xie, Zhihui Li, Bing Wang, Xiaojun Chang

专题命中 空间理解 :novel view synthesis(abstract);分类 cs.CV

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.13199 2025-03-11 cs.CV 57%

MGNet: Monocular Geometric Scene Understanding for Autonomous Driving

Markus Schön, Michael Buchholz, Klaus Dietmayer

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 15784-15795

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05596 2025-02-25 cs.CV 57%

TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances

Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00255 2025-02-21 cs.AI cs.CL cs.CV 57%

Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Weitai Kang, Haifeng Huang, Yuzhang Shang, Mubarak Shah, Yan Yan

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏