arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

3D 视觉

三维重建、NeRF、Gaussian Splatting、点云和空间智能。

共收录 497 信号源:cs.CV, cs.GR, cs.RO

1. 空间理解 497 篇

2602.13588 2026-02-17 cs.CV cs.AI 57%

Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks

双流交互式场景解析与几何视觉任务联合学习

Guanfeng Tang, Hongbo Zhao, Ziwei Long, Jiayao Li, Bohong Xiao, Wei Ye, Hanli Wang, Rui Fan

机构 * College of Electronic and Information Engineering, Tongji University(同济大学电子信息学院) Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统研究所) College of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) Department of Vehicle Control System and Software Development, NIO(蔚来汽车车辆控制系统与软件开发部门) Department of Automotive Engineering, Jilin University(吉林大学汽车工程学院) Key Laboratory of Embedded System and Service Computing (Ministry of Education), Tongji University(同济大学嵌入式系统与服务计算重点实验室) College of Electronic and Information Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院电子信息学院) Shanghai Key Laboratory of Intelligent Autonomous Systems(上海智能自主系统重点实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 TwInS通过双流交互式联合学习框架,实现场景解析与几何视觉任务的同步优化,采用跨任务适配器提升性能并减少对人工标注的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21769 2026-02-12 cs.CV 57%

H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows

H2OFlow: 通过3D生成模型和密集扩散流接地人类-物体 affordances

Harry Zhang, Luca Carlone

机构 * MIT(麻省理工学院)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 H2OFlow通过3D生成模型和密集扩散流学习人类-物体交互的3D affordances,无需人工标注,有效泛化至现实物体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 57%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21602 2026-02-04 cs.RO 57%

AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation

AIR-VLA:面向空中操作的视觉-语言-动作系统

Jianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian

机构 * The Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) School of Information and Intelligent Engineering, University of Sanya(三亚大学信息与智能工程学院) School of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) State Key Lab of Intelligent Transportation Systems, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

AI总结 AIR-VLA提出首个针对空中操作的视觉-语言-动作系统,通过构建仿真环境和多模态数据集,评估主流模型并揭示其在无人机移动、机械臂控制和高层规划中的能力和限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21634 2026-01-30 cs.CV 57%

RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning

RSGround-R1: 重新思考通过空间推理的遥感视觉定位

Shiqi Huang, Shuting He, Bihan Wen

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(电气电子工程学院,南洋理工大学) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(教育部计算与经济交叉学科重点实验室,上海财经大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 RSGround-R1通过引入空间推理引导的后训练框架,提升遥感图像中目标物体的定位精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11574 2026-01-19 cs.CV 57%

Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis

通过多视角对应分析评估基础模型的3D理解

Valentina Lilova, Toyesh Chakravorty, Julian I. Bibo, Emma Boccaletti, Brandon Li, Lívia Baxová, Cees G. M. Snoek, Mohammadreza Salehi

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出一种无需微调的3D场景理解基准,评估基础模型在多视角下的3D推理能力,展示DINO编码器在大视角变化下的竞争力。

Comments NeurIPS 2025 UniReps workshop, to be published in PMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08728 2026-01-14 cs.CV 57%

Salience-SGG: Enhancing Unbiased Scene Graph Generation with Iterative Salience Estimation

显著性场景图生成:通过迭代显著性估计增强无偏场景图生成

Runfeng Qu, Ole Hall, Pia K Bideau, Julie Ouerfelli-Ethier, Martin Rolfs, Klaus Obermayer, Olaf Hellwich

机构 * Technische Universität Berlin(柏林技术大学) Humboldt Universität zu Berlin(柏林洪堡大学) Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、法国国家信息与自动化研究所、法国国家科学研究中心、格勒诺布尔INP、LJK) Bernstein Center for Computational Neuroscience(计算神经科学伯恩斯坦中心) Science of Intelligence Research Cluster of Excellence(卓越人工智能研究集群)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 Salience-SGG通过迭代显著性估计提升场景图生成的无偏性,改进了空间理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22976 2026-01-13 cs.CV 57%

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

从二维空间到三维空间:教视觉语言模型感知和推理三维世界

Jiahui Zhang, Yurui Chen, Yanpeng Zhou, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yu-Jie Yuan, Xinyue Cai, Guowei Huang, Xingyue Quan, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出了一种基于2D空间数据生成和标注的管道,构建了SPAR-7M数据集和SPAR-Bench基准,通过训练和微调提升视觉语言模型在三维空间感知和推理中的性能。

Comments Project page: https://logosroboticsgroup.github.io/SPAR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06781 2026-01-13 cs.HC cs.AI cs.CV 57%

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

AutoTour:基于智能手机和LLMs的自动照片导览系统

Huatao Xu, Zihe Liu, Zilin Zeng, Baichuan Li, Mo Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 AutoTour利用智能手机和LLMs自动为用户照片生成细粒度地标注释和描述,实现可扩展且上下文感知的交互式探索体验。

Comments 21

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03590 2026-01-08 cs.CV cs.AI 57%

Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions

大语言模型能否通过像素感知空间智能?基于文本描述的空间智能基准测试

Zhongbin Guo, Zhen Yang, Yushan Li, Xinyue Zhang, Wenyu Gao, Jiacheng Wang, Chengzhi Li, Xiangrui Liu, Ping Jian

机构 * School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 SiT-Bench通过文本描述评估大语言模型的空间智能,揭示其在全局一致性方面的不足,并表明显式空间推理可提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13683 2026-01-08 cs.CV 57%

I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners

I-Scene: 3D实例模型是隐式可泛化的空间学习者

Lu Ling, Yunhao Ge, Yichen Sheng, Aniket Bera

机构 * Purdue University(普渡大学) NVIDIA Research(英伟达研究)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 I-Scene提出了一种基于3D实例模型的隐式空间学习方法,通过重新编程生成器实现跨场景的泛化能力,展示了其在空间推理和生成中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01416 2026-01-06 cs.CV 57%

AirSpatialBot: A Spatially-Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognization and Retrieval

AirSpatialBot: 一种空间感知的空中智能体用于细粒度车辆属性识别与检索

Yue Zhou, Ran Ding, Xue Yang, Xue Jiang, Xingzhao Liu

机构 * Department of Electronic Engineering, Shanghai Jiao Tong University(电子工程系,上海交通大学) Department of Automation, Shanghai Jiao Tong University(自动化系,上海交通大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 AirSpatialBot通过空间感知的VLM和两阶段训练策略,实现细粒度车辆属性识别与检索,解决遥感图像中的空间理解难题。

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01850 2026-01-06 cs.LG cs.RO 57%

ManiBox: Enhancing Embodied Spatial Generalization via Scalable Simulation Data Generations

ManiBox: 通过可扩展的模拟数据生成增强具身空间泛化

Hengkai Tan, Xuezhou Xu, Chengyang Ying, Xinyi Mao, Zeyuan Wang, Songming Liu, Xingxing Zhang, Zhizhong Su, Hang Su, Jun Zhu

机构 * Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Horizon Robotics

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

AI总结 ManiBox通过可扩展的模拟数据生成提升具身智能体的空间泛化能力,有效减少仿真到现实的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06259 2025-12-29 cs.CV cs.AI 57%

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

SIFThinker: 基于空间的图像聚焦用于视觉推理

Zhangquan Chen, Ruihui Zhao, Chuwei Luo, Mingze Sun, Xinlei Yu, Yangyang Kang, Ruqi Huang

机构 * ByteDance(字节跳动)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 SIFThinker通过结合深度增强的边界框和自然语言,提升视觉推理中的空间理解和细粒度感知能力。

Comments 15 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08712 2025-12-25 cs.RO 57%

NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance

NavDP: 基于特权信息引导的学习导航扩散策略

Wenzhe Cai, Jiaqi Peng, Yuqiang Yang, Yujian Zhang, Meng Wei, Hanqing Wang, Yilun Chen, Tai Wang, Jiangmiao Pang

机构 * Shanghai AI Lab(上海人工智能实验室) Tsinghua University(清华大学) Zhejiang University(浙江大学) The University of Hong Kong(香港大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

AI总结 NavDP通过基于特权信息引导的学习方法,实现自主机器人在多样化环境中的零样本模拟到现实导航转移。

Comments Project Page: https://wzcai99.github.io/navigation-diffusion-policy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20501 2025-12-24 cs.CV 57%

Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition

弥合模态与知识转移:增强多模态理解和识别

Gorjan Radevski

机构 * University of Bristol(布里斯托大学) University of Würzburg(乌尔姆大学) Processing Speech and Images (PSI)(语音与图像处理组)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出多模态对齐、翻译、融合和转移方法,提升复杂输入的理解与识别能力,涵盖空间语言、医学文本、知识图谱和动作识别等多个领域。

Comments Ph.D. manuscript; Supervisors/Mentors: Marie-Francine Moens and Tinne Tuytelaars

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13387 2025-12-24 cs.CV eess.IV 57%

From Binary to Semantic: Utilizing Large-Scale Binary Occupancy Data for 3D Semantic Occupancy Prediction

从二元到语义:利用大规模二元占用数据进行3D语义占用预测

Chihiro Noguchi, Takaki Yamamoto

机构 * InfoTech, Toyota Motor Corporation(丰田汽车公司信息科技部)

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 本文提出一种基于二元占用数据的框架,通过预训练和自动标注方法提升3D语义占用预测的性能。

Comments Accepted to ICCV Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19609 2025-12-23 cs.CV cs.AI 57%

MapTrace: Scalable Data Generation for Route Tracing on Maps

MapTrace: 用于地图路线追踪的可扩展数据生成

Artemis Panagopoulou, Aveek Purohit, Achin Kulshrestha, Soroosh Yazdani, Mohit Goyal

机构 * Google XR(谷歌XR) University of Pennsylvania(宾夕法尼亚大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 MapTrace通过合成数据生成提升地图路线追踪性能,微调模型使成功率提升6.4个百分点,减少路径追踪误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14020 2025-12-17 cs.CV 57%

Deep Learning Perspective of Scene Understanding in Autonomous Robots

自主机器人场景理解的深度学习视角

Afia Maham, Dur E Nayab Tashfa

专题命中 空间理解 :3D reconstruction(abstract);分类 cs.CV

AI总结 本文从深度学习角度综述了自主机器人场景理解中的关键技术,包括目标检测、语义分割、深度估计等,探讨了其在动态环境中的应用与挑战。

Comments 11 pages. Review Paper on Deep Learning Perspective of Scene Understanding in Autonomous Robots

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03682 2025-12-16 cs.CV cs.AI cs.LG 57%

How PARTs assemble into wholes: Learning the relative composition of images

如何将PART组合成整体:学习图像的相对组成

Melika Ayoughi, Samira Abnar, Chen Huang, Chris Sandino, Sayeri Lala, Eeshan Gunesh Dhekane, Dan Busbridge, Shuangfei Zhai, Vimal Thilak, Josh Susskind, Pascal Mettes, Paul Groth, Hanlin Goh

机构 * University of Amsterdam(阿姆斯特丹大学) Apple(苹果公司)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 PART通过连续相对变换学习图像的相对组成,优于基于网格的方法,在空间理解任务中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07472 2025-12-09 cs.RO cs.LG 57%

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

可及场干预:使VLAs在机器人操作中摆脱记忆陷阱

Siyu Xu, Zijian Wang, Yunke Wang, Chenghao Xia, Tao Huang, Chang Xu

机构 * School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) John Hopcropt Center for Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学中心)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

AI总结 本研究提出可及场干预(AFI)方法,通过引入3D空间可及场提升VLA在机器人操作中对分布变化的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04686 2025-12-09 cs.CV 57%

Towards Cross-View Point Correspondence in Vision-Language Models

面向视觉-语言模型的跨视图点对应研究

Yipu Wang, Yuheng Ji, Yuyang Liu, Enshen Zhou, Ziqiang Yang, Yuxuan Tian, Ziheng Qin, Yue Liu, Huajie Tan, Cheng Chi, Zhiyuan Ma, Daniel Dajun Zeng, Xiaolong Zheng

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Beihang University(北航大学) Jilin University(吉林大学) National University of Singapore(新加坡国立大学) Peking University(北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Huazhong University of Science And Technology(华中科技大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 本文提出跨视图点对应任务和 CrossPoint-Bench 基准,通过 CroPond 模型在 CrossPoint-Bench 上取得超越 Gemini-2.5-Pro 的准确率,推动跨视图对应研究发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13558 2025-12-09 cs.CV 57%

X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability

X-Scene: 通过高保真和灵活可控性实现大规模驾驶场景生成

Yu Yang, Alan Liang, Jianbiao Mei, Yukai Ma, Yong Liu, Gim Hee Lee

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 空间理解 :3DGS(abstract);分类 cs.CV

AI总结 X-Scene通过高保真和灵活可控性实现大规模驾驶场景生成,支持多粒度控制和一致性外推,提升自动驾驶的数据生成与模拟能力。

Comments Accepted by NeurIPS 2025, Project page at https://x-scene.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23075 2025-12-05 cs.CV cs.AI 57%

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

SpaceMind: 基于摄像头引导的模态融合用于视觉-语言模型中的空间推理

Ruosen Zhao, Zhikang Zhang, Jialei Xu, Jiahao Chang, Dong Chen, Lingyun Li, Weijian Sun, Zizhuang Wei

机构 * Huawei(华为) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) The University of Hong Kong(香港大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 SpaceMind通过摄像头引导的模态融合方法,提升视觉-语言模型在空间推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02715 2025-12-03 cs.CV 57%

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

GeoViS: 基于地理奖励的视觉搜索用于遥感视觉定位

Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxiang Wang, Xue Yang, Kaiwen Wei, Lei Wang

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) National University of Defense Technology(国防科技大学) Shanghai Jiao Tong University(上海交通大学) Chongqing University(重庆大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 GeoViS通过地理奖励机制,实现了遥感影像中的细粒度视觉定位,通过逐步搜索和推理提升小目标检测精度与跨领域泛化能力。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03342 2025-12-02 cs.RO 57%

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Gemini Robotics 1.5:通过先进的具身推理、思考和运动转移推动通用机器人前沿

Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Ashwin Balakrishna, Nathan Batchelor, Alex Bewley, Jeff Bingham, Michael Bloesch, Konstantinos Bousmalis, Philemon Brakel, Anthony Brohan, Thomas Buschmann, Arunkumar Byravan, Serkan Cabi, Ken Caluwaerts, Federico Casarini, Christine Chan, Oscar Chang, London Chappellet-Volpini, Jose Enrique Chen, Xi Chen, Hao-Tien Lewis Chiang, Krzysztof Choromanski, Adrian Collister, David B. D'Ambrosio, Sudeep Dasari, Todor Davchev, Meet Kirankumar Dave, Coline Devin, Norman Di Palo, Tianli Ding, Carl Doersch, Adil Dostmohamed, Yilun Du, Debidatta Dwibedi, Sathish Thoppay Egambaram, Michael Elabd, Tom Erez, Xiaolin Fang, Claudio Fantacci, Cody Fong, Erik Frey, Chuyuan Fu, Ruiqi Gao, Marissa Giustina, Keerthana Gopalakrishnan, Laura Graesser, Oliver Groth, Agrim Gupta, Roland Hafner, Steven Hansen, Leonard Hasenclever, Sam Haves, Nicolas Heess, Brandon Hernaez, Alex Hofer, Jasmine Hsu, Lu Huang, Sandy H. Huang, Atil Iscen, Mithun George Jacob, Deepali Jain, Sally Jesmonth, Abhishek Jindal, Ryan Julian, Dmitry Kalashnikov, M. Emre Karagozler, Stefani Karp, Matija Kecman, J. Chase Kew, Donnie Kim, Frank Kim, Junkyung Kim, Thomas Kipf, Sean Kirmani, Ksenia Konyushkova, Li Yang Ku, Yuheng Kuang, Thomas Lampe, Antoine Laurens, Tuan Anh Le, Isabel Leal, Alex X. Lee, Tsang-Wei Edward Lee, Guy Lever, Jacky Liang, Li-Heng Lin, Fangchen Liu, Shangbang Long, Caden Lu, Sharath Maddineni, Anirudha Majumdar, Kevis-Kokitsi Maninis, Andrew Marmon, Sergio Martinez, Assaf Hurwitz Michaely, Niko Milonopoulos, Joss Moore, Robert Moreno, Michael Neunert, Francesco Nori, Joy Ortiz, Kenneth Oslund, Carolina Parada, Emilio Parisotto, Amaris Paryag, Acorn Pooley, Thomas Power, Alessio Quaglino, Haroon Qureshi, Rajkumar Vasudeva Raju, Helen Ran, Dushyant Rao, Kanishka Rao, Isaac Reid, David Rendleman, Krista Reymann, Miguel Rivas, Francesco Romano, Yulia Rubanova, Peter Pastor Sampedro, Pannag R Sanketi, Dhruv Shah, Mohit Sharma, Kathryn Shea, Mohit Shridhar, Charles Shu, Vikas Sindhwani, Sumeet Singh, Radu Soricut, Rachel Sterneck, Ian Storz, Razvan Surdulescu, Jie Tan, Jonathan Tompson, Saran Tunyasuvunakool, Jake Varley, Grace Vesom, Giulia Vezzani, Maria Bauza Villalonga, Oriol Vinyals, René Wagner, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Chengda Wu, Markus Wulfmeier, Fei Xia, Ted Xiao, Annie Xie, Jinyu Xie, Peng Xu, Sichun Xu, Ying Xu, Zhuo Xu, Jimmy Yan, Sherry Yang, Skye Yang, Yuxiang Yang, Hiu Hong Yu, Wenhao Yu, Wentao Yuan, Yuan Yuan, Jingwei Zhang, Tingnan Zhang, Zhiyuan Zhang, Allan Zhou, Guangyao Zhou, Yuxiang Zhou

专题命中 空间理解 :spatial understanding(abstract);分类 cs.RO

AI总结 Gemini Robotics 1.5通过具身推理、思考和运动转移技术,提升通用机器人在复杂任务中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12718 2025-11-27 cs.CV 57%

EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer

EvoEmpirBench: 基于智能体经验的动态空间推理

Pukun Zhao, Longxiang Wang, Miaowei Wang, Chen Chen, Fanqing Zhou, Haojian Huang

机构 * Guangdong University of Finance and Economics(广东金融学院) Chongqing University(重庆大学) University of Edinburgh(爱丁堡大学) The University of Hong Kong(香港大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 EvoEmpirBench通过动态空间基准评估模型在部分可观测和动态环境下的空间推理与记忆能力,揭示主流模型的限制并提供未来研究平台。

Comments Accepted by AAAI 2026, 29 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19261 2025-11-25 cs.CV 57%

LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models

LAST: 为通用视觉-语言模型学习在空间和时间中的思考

Shuai Wang, Daoan Zhang, Tianyi Bai, Shitong Shao, Jiebo Luo, Jiaheng Wei

机构 * HKUST(GZ)(香港科技大学(广州)) University of Rochester(罗切斯特大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 LAST通过学习在空间和时间上思考,提升通用视觉-语言模型对3D空间和长视频的理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17940 2025-11-25 cs.CV 57%

Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames

场景摘要:将场景视频聚类为空间上多样的帧

Chao Chen, Mingzhi Zhu, Ankush Pratap Singh, Yu Yan, Felix Juefei-Xu, Chen Feng

机构 * New York University(纽约大学)

专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV

AI总结 SceneSum通过自监督方法将场景视频聚类为空间多样化的关键帧,提升空间推理能力,优于现有视频摘要方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17048 2025-11-24 cs.CV 57%

RoomPlanner: Explicit Layout Planner for Easier LLM-Driven 3D Room Generation

RoomPlanner: 一种显式布局规划器,用于更易由LLM驱动的3D房间生成

Wenzhuo Sun, Mingjian Liang, Wenxuan Song, Xuelian Cheng, Zongyuan Ge

专题命中 空间理解 :point cloud(abstract);分类 cs.CV

AI总结 RoomPlanner通过显式布局规划和高效优化策略,实现了快速生成高质量3D室内场景。

详情

展开后加载摘要…

URL PDF HTML 收藏