arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3076 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3076 篇

2605.11367 2026-06-01 cs.CV 85%

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

3D-Belief:通过生成式3D世界建模实现具身信念推断

Yifan Yin, Zehao Wen, Suyu Ye, Jieneng Chen, Zehan Zheng, Nanru Dai, Haojun Shi, Aydan Huang, Zheyuan Zhang, Alan Yuille, Jianwen Xie, Ayush Tewari, Tianmin Shu

机构 * Johns Hopkins University(约翰霍普金斯大学) Lambda University of Cambridge(剑桥大学)

专题命中 具身推理 :world model(title,abstract);embodied agent(abstract);navigation(abstract);分类 cs.CV

AI总结 提出3D-Belief,一种生成式3D世界模型,通过在线更新显式3D信念,使具身智能体能够在部分可观测环境中想象场景补全并推理,在2D/3D想象质量和下游物体导航任务上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11351 2026-04-14 cs.RO 85%

WM-DAgger: Enabling Efficient Data Aggregation for Imitation Learning with World Models

WM-DAgger: 通过世界模型实现模仿学习中的高效数据聚合

Anlan Yu, Zaishu Chen, Peili Song, Zhiqing Hong, Haotian Wang, Desheng Zhang, Tian He, Yi Ding, Daqing Zhang

机构 * Peking University(北京大学) JD Logistics(京东物流) Nankai University(南开大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Rutgers University(罗格斯大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) Institut Polytechnique de Paris(巴黎综合理工学院)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出WM-DAgger框架,利用世界模型合成OOD恢复数据,无需人工干预,提升机器人模仿学习效率,实验表明在少样本下显著提高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12639 2026-04-14 cs.CV 85%

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

RoboStereo: 双塔4D具身世界模型用于统一策略优化

Ruicheng Zhang, Guangyu Chen, Zunnan Xu, Zihao Liu, Zhizhou Zhong, Mingyang Zhang, Jun Zhou, Xiu Li

机构 * Tsinghua University(清华大学) X Square Robot HKUST(香港科技大学)

专题命中 具身推理 :world model(title,abstract);embodied AI(abstract);manipulation(abstract);分类 cs.CV

AI总结 本文提出RoboStereo,通过双塔4D具身世界模型实现统一策略优化,解决几何幻觉和缺乏统一优化框架的问题,实验显示在细粒度操作任务中平均相对提升超过97%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21714 2026-04-09 cs.CV 85%

AstraNav-World: World Model for Foresight Control and Consistency

AstraNav-World:面向前瞻性控制与一致性的世界模型

Jintao Chen, Junjun Hu, Haochen Bai, Minghua Luo, Xinda Xue, Botao Ren, Chengyu Bai, Shichao Xie, Ziyi Chen, Fei Liu, Zedong Chu, Xiaolong Wu, Mu Xu, Shanghang Zhang

机构 * Amap Alibaba(高德阿里巴巴)

专题命中 具身推理 :world model(title,abstract);embodied agent(abstract);navigation(abstract);分类 cs.CV

AI总结 本文提出AstraNav-World,通过统一的概率框架联合推理未来视觉状态与动作序列,提升开放动态环境中的导航轨迹精度与成功率,实验表明其在真实场景中具备零样本适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16860 2026-03-18 cs.RO 85%

DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models

DreamPlan: 通过视频世界模型高效强化微调视觉语言规划器

Emily Yue-Ting Jia, Weiduo Yuan, Tianheng Shi, Vitor Guizilini, Jiageng Mao, Yue Wang

机构 * USC Physical Superintelligence Lab(USC物理超智能实验室) Toyota Research Institute(丰田研究院)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 DreamPlan通过视频世界模型高效强化微调视觉语言规划器,利用零样本VLM生成探索数据训练动作条件视频生成模型,再通过ORPO在虚拟环境中微调VLM,提升物体 manipulation 成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01960 2026-03-17 cs.LG 85%

Grounding Generated Videos in Feasible Plans via World Models

通过世界模型将生成视频接地到可行计划中

Christos Ziakas, Amir Bar, Alessandra Russo

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);navigation(abstract);分类 cs.LG

AI总结 本文提出GVP-WM方法,通过学习动作条件的世界模型,将生成视频计划接地为可行动作序列,解决视频生成计划在时间一致性和物理约束上的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00296 2026-03-10 cs.RO cs.AI cs.CV cs.LG 85%

From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models

从像素到谓词:通过预训练视觉-语言模型学习符号世界模型

Ashay Athalye, Nishanth Kumar, Tom Silver, Yichao Liang, Jiuguang Wang, Tomás Lozano-Pérez, Leslie Pack Kaelbling

机构 * MIT(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) RAI Institute(RAI研究院)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 通过预训练视觉-语言模型学习符号世界模型,以实现复杂机器人领域中长周期决策制定的零样本泛化。

Comments A version of this paper appears in the official proceedings of RA-L, Volume 11, Issue 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19400 2026-03-03 cs.CV 85%

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

跨视角视觉:评估视觉-语言模型在机器人场景中的空间推理能力

Zhiyuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du, Jiongrui Yan, Shubin Shi, Chengbo Yuan, Huizhi Liang, Yu Deng, Qixiu Li, Rushuai Yang, Arctanx An, Leqi Zheng, Weijie Wang, Shawn Chen, Sicheng Xu, Yaobo Liang, Jiaolong Yang, Baining Guo

机构 * Tsinghua University(清华大学) Peking University(北京大学) Fudan University(复旦大学) Microsoft Research Asia(微软亚洲研究院) Hong Kong University of Science and Technology(香港科技大学) Zhejiang University(浙江大学)

专题命中 具身推理 :robotic(title,abstract);embodied AI(abstract);manipulation(abstract);分类 cs.CV

AI总结 本文提出MV-RoboBench基准,评估视觉-语言模型在机器人场景中的多视角空间推理能力,揭示其在多视角机器人感知中的挑战。

Comments Accepted to ICLR 2026. Camera-ready version. Project page: https://aaronfengzy.github.io/MV-RoboBench-Webpage/

Journal ref International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12099 2026-02-27 cs.CV 85%

GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

GigaBrain-0.5M*: 基于世界模型强化学习的VLA

GigaBrain Team, Boyuan Wang, Bohan Li, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Jie Li, Jindi Lv, Jingyu Liu, Lv Feng, Mingming Yu, Peng Li, Qiuping Deng, Tianze Liu, Xinyu Zhou, Xinze Chen, Xiaofeng Wang, Yang Wang, Yifan Li, Yifei Nie, Yilong Li, Yukun Zhou, Yun Ye, Zhichao Liu, Zheng Zhu

机构 * GigaBrain Team(GigaBrain团队)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.CV

AI总结 GigaBrain-0.5M*通过基于世界模型的强化学习提升VLA模型的跨任务适应能力,实现复杂操作任务的可靠执行。

Comments https://gigabrain05m.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06949 2026-02-09 cs.RO cs.AI cs.CV cs.LG 85%

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

DreamDojo:从大规模人类视频中学习通用机器人世界模型

Shenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik, Seonghyeon Ye, Sihyun Yu, Wei-Cheng Tseng, Yuzhu Dong, Kaichun Mo, Chen-Hsuan Lin, Qianli Ma, Seungjun Nah, Loic Magne, Jiannan Xiang, Yuqi Xie, Ruijie Zheng, Dantong Niu, You Liang Tan, K. R. Zentner, George Kurian, Suneel Indupuru, Pooya Jannaty, Jinwei Gu, Jun Zhang, Jitendra Malik, Pieter Abbeel, Ming-Yu Liu, Yuke Zhu, Joel Jang, Linxi "Jim" Fan

机构 * NVIDIA HKUST(香港科技大学) UC Berkeley(加州大学伯克利分校) Stanford(斯坦福大学) KAIST(韩国科学技术院) UofT(多伦多大学) UCSD(加州大学圣地亚哥分校) UT Austin(德克萨斯大学奥斯汀分校)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 DreamDojo通过大规模人类视频学习通用机器人世界模型,解决动作标签稀缺问题,实现高精度物理理解和实时交互。

Comments Project page: https://dreamdojo-world.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03077 2025-11-06 cs.RO 85%

WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models

R. Khorrambakht, Joaquim Ortiz-Haro, Joseph Amigo, Omar Mostafa, Daniel Dugas, Franziska Meier, Ludovic Righetti

机构 * Center for Robotics and Embodied Intelligence, Tandon School of Engineering, New York University, Brooklyn, NY(机器人与具身智能中心,工程学院,纽约大学,布鲁克林,纽约) FAIR at Meta(Meta的FAIR)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21657 2025-11-03 cs.CV 85%

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

Yixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu, Yonggang Qi

机构 * AMAP, Alibaba Group(阿里集团AMAP) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 具身推理 :world model(title,abstract);navigation(abstract);robotic(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10912 2025-10-16 cs.RO 85%

More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks

Xinyu Shao, Yanzhe Tang, Pengwei Xie, Kaiwen Zhou, Yuzheng Zhuang, Xingyue Quan, Jianye Hao, Long Zeng, Xiu Li

机构 * Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院) Noah’s Ark Lab, Huawei, China(华为诺亚实验室)

专题命中 具身推理 :robotic(title,abstract);manipulation(abstract);navigation(abstract);分类 cs.RO

Comments More details and videos can be found at https://robo-map.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21790 2025-09-29 cs.CV 85%

LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE

Yu Shang, Lei Jin, Yiding Ma, Xin Zhang, Chen Gao, Wei Wu, Yong Li

机构 * Tsinghua University(清华大学) Manifold AI

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.CV

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05217 2025-06-06 cs.CV 85%

DSG-World: Learning a 3D Gaussian World Model from Dual State Videos

Wenhao Hu, Xuexiang Wen, Xi Li, Gaoang Wang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ZJU-UIUC Institute, Zhejiang University(浙江大学-伊利诺伊大学联合研究所)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);manipulation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21232 2025-03-28 cs.AI 85%

Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles

Ayush Bheemaiah, Seungyong Yang

专题命中 具身推理 :world model(title,abstract);robotics(abstract);embodied AI(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23156 2025-03-04 cs.AI cs.CV cs.LG cs.RO 85%

VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning

Yichao Liang, Nishanth Kumar, Hao Tang, Adrian Weller, Joshua B. Tenenbaum, Tom Silver, João F. Henriques, Kevin Ellis

专题命中 具身推理 :world model(title,abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.CV

Comments ICLR 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08204 2025-01-24 cs.RO 85%

SAILOR: Perceptual Anchoring For Robotic Cognitive Architectures

Miguel Á. González-Santamarta, Francisco J. Rodríguez-Lera, Vicente Matellán Olivera, Virginia Riego Del Castillo, Lidia Sánchez-González

专题命中 具身推理 :robotic(title,abstract);robotics(abstract);world model(abstract);分类 cs.RO

Comments 16 pages, 5 figures, 7 tables, 3 algorithms, Submitted to Scientific Reports

Journal ref Scientific Reports 15, 113 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11870 2024-09-19 cs.RO 85%

SpotLight: Robotic Scene Understanding through Interaction and Affordance Detection

Tim Engelbracht, René Zurbrügg, Marc Pollefeys, Hermann Blum, Zuria Bauer

专题命中 具身推理 :robotic(title,abstract);robotics(abstract);robot learning(abstract);分类 cs.RO

Comments timengelbracht.github.io/SpotLight/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11289 2024-08-23 cs.RO 85%

ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models

Siyuan Huang, Iaroslav Ponomarenko, Zhengkai Jiang, Xiaoqi Li, Xiaobin Hu, Peng Gao, Hongsheng Li, Hao Dong

专题命中 具身推理 :robotic(title,abstract);robotics(abstract);manipulation(abstract);分类 cs.RO

Comments Code and dataset are publicly available at https://github.com/SiyuanHuang95/ManipVQA. Accepted by IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09056 2024-04-04 cs.LG cs.AI cs.CV cs.RO stat.ML 85%

ReCoRe: Regularized Contrastive Representation Learning of World Model

Rudra P. K. Poudel, Harit Pandya, Stephan Liwicki, Roberto Cipolla

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.CV

Comments Accepted at CVPR 2024. arXiv admin note: text overlap with arXiv:2209.14932

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17593 2023-11-30 cs.LG cs.AI cs.CL cs.CV cs.RO 85%

LanGWM: Language Grounded World Model

Rudra P. K. Poudel, Harit Pandya, Chao Zhang, Roberto Cipolla

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10901 2023-08-22 cs.RO cs.AI cs.CV cs.LG cs.NE 85%

Structured World Models from Human Videos

Russell Mendonca, Shikhar Bahl, Deepak Pathak

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

Comments RSS 2023. Website at https://human-world-model.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14932 2022-09-30 cs.LG cs.AI cs.CV cs.RO stat.ML 85%

Contrastive Unsupervised Learning of World Model with Invariant Causal Features

Rudra P. K. Poudel, Harit Pandya, Roberto Cipolla

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13022 2022-04-28 cs.LG 85%

Binding Actions to Objects in World Models

Ondrej Biza, Robert Platt, Jan-Willem van de Meent, Lawson L. S. Wong, Thomas Kipf

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.LG

Comments Published at the ICLR 2022 workshop on Objects, Structure and Causality

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10335 2022-01-26 cs.LG 85%

Tracking and Planning with Spatial World Models

Baris Kayalibay, Atanas Mirchev, Patrick van der Smagt, Justin Bayer

专题命中 具身推理 :world model(title,abstract);robotics(abstract);navigation(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22796 2026-03-25 cs.CV cs.AI cs.RO 85%

PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding

PhotoAgent:一种具有空间与审美理解的机器人摄影师

Lirong Che, Zhenfeng Gan, Yanbo Chen, Junbo Tan, Xueqian Wang

机构 * Center for Artificial Intelligence and Robotics, Shenzhen International Graduate School, Tsinghua University(人工智能与机器人中心,深圳国际研究生院,清华大学)

专题命中 具身推理 :robotic(title);embodied agent(abstract);world model(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 PhotoAgent通过整合大多模态模型与新型控制范式,将主观审美目标转化为几何约束,利用内部视觉模拟实现快速收敛于高质量图像结果。

Comments Accepted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01924 2025-12-02 cs.RO cs.AI cs.LG 85%

Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model

通过时序分层世界模型的深度主动推断实现现实世界的机器人控制

Kentaro Fujii, Shingo Murata

机构 * Graduate School of Integrated Design Engineering, Keio University(Keio大学整合设计工程研究院)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.LG;robotics(comments)

AI总结 本文提出一种结合时序分层世界模型的深度主动推断框架,用于在不确定环境中实现机器人高成功率的操作与探索性动作切换。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07092 2025-10-09 cs.LG cs.AI cs.RO 85%

Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report

Riccardo Mereu, Aidan Scannell, Yuxin Hou, Yi Zhao, Aditya Jitta, Antonio Dominguez, Luigi Acerbi, Amos Storkey, Paul Chang

机构 * Aalto University(阿alto大学) University of Edinburgh(爱丁堡大学) Deep Render(Deep Render公司) DataCrunch(DataCrunch公司) University of Helsinki(赫尔辛基大学)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);分类 cs.RO、cs.AI、cs.LG

Comments 6 pages, 3 figures, 1X world model challenge technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18926 2024-04-30 cs.RO cs.CV cs.LG 85%

Point Cloud Models Improve Visual Robustness in Robotic Learners

Skand Peri, Iain Lee, Chanho Kim, Li Fuxin, Tucker Hermans, Stefan Lee

专题命中 具身推理 :robotic(title,abstract);world model(abstract);分类 cs.RO、cs.CV、cs.LG;robotics(comments)

Comments Accepted at International Conference on Robotics and Automation, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏