arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-03-10 至 2026-03-10 共收录 14 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 14 篇

2505.14357 2026-03-10 cs.CV cs.LG 86%

Vid2World: Crafting Video Diffusion Models to Interactive World Models

Vid2World: 构建视频扩散模型以交互式世界模型

Siqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao, Mingsheng Long

机构 * Tsinghua University(清华大学) Chongqing University(重庆大学)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);navigation(abstract);分类 cs.CV、cs.LG

AI总结 Vid2World通过因果化视频扩散模型,提升其在交互式世界建模中的可控性和生成能力,适用于机器人操作、游戏模拟和开放世界导航等多种领域。

Comments Project page: http://knightnemo.github.io/vid2world/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00296 2026-03-10 cs.RO cs.AI cs.CV cs.LG 85%

From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models

从像素到谓词:通过预训练视觉-语言模型学习符号世界模型

Ashay Athalye, Nishanth Kumar, Tom Silver, Yichao Liang, Jiuguang Wang, Tomás Lozano-Pérez, Leslie Pack Kaelbling

机构 * MIT(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) RAI Institute(RAI研究院)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 通过预训练视觉-语言模型学习符号世界模型,以实现复杂机器人领域中长周期决策制定的零样本泛化。

Comments A version of this paper appears in the official proceedings of RA-L, Volume 11, Issue 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07799 2026-03-10 cs.CV cs.RO 84%

MWM: Mobile World Models for Action-Conditioned Consistent Prediction

MWM: 移动世界模型用于动作条件一致预测

Han Yan, Zishang Xiang, Zeyu Zhang, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.RO、cs.CV

AI总结 MWM提出了一种移动世界模型,通过两阶段训练框架和推理一致状态蒸馏提升动作条件一致性的图像目标导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05522 2026-03-10 cs.AI cs.CV cs.LG cs.RO 83%

RoboLayout: Differentiable 3D Scene Generation for Embodied Agents

RoboLayout: 用于具身智能体的可微3D场景生成

Ali Shamsaddinlou

专题命中 具身推理 :embodied agent(title,abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 RoboLayout通过引入智能体感知推理和改进优化稳定性,提升3D场景生成在具身智能体交互中的可行性与物理合理性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08455 2026-03-10 cs.AI cs.LG 81%

The Boiling Frog Threshold: Criticality and Blindness in World Model-Based Anomaly Detection Under Gradual Drift

青蛙沸腾阈值:世界模型基于异常检测中的临界性与盲目性

Zhe Hong

机构 * National University of Singapore(新加坡国立大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 研究揭示了世界模型在异常检测中的临界阈值,发现检测阈值受噪声底座、检测器和环境动态的三重交互影响,且正弦漂移无法被检测到。

Comments 10 pages, 5 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11682 2026-03-10 cs.RO cs.AI cs.SY eess.SY 81%

Ego-Vision World Model for Humanoid Contact Planning

人形机器人接触规划的视角世界模型

Hang Liu, Yuman Gao, Sangli Teng, Yufeng Chi, Yakun Sophia Shao, Zhongyu Li, Maani Ghaffari, Koushil Sreenath

机构 * University of California, Berkeley(加州大学伯克利分校) University of Michigan, Ann Arbor(密歇根大学安娜堡分校) The Chinese University of Hong Kong(香港中文大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种结合学习世界模型与MPC的框架,用于提升人形机器人在复杂环境中的接触规划能力,实现更高效和稳健的多任务执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07264 2026-03-10 cs.RO cs.AI 81%

Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving

具备运动学意识的潜在世界模型用于数据高效的自动驾驶

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

机构 * Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University(道路与交通工程重点实验室,教育部,同济大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出一种具备运动学意识的潜在世界模型框架,通过整合车辆运动学信息提升自动驾驶的样本效率和驾驶性能。

Comments 6 pages, 5 figures. Under review at IEEE ITSC

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19195 2026-03-10 cs.CV cs.AI 81%

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

重新思考驾驶世界模型作为感知任务的合成数据生成器

Kai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei, Xiangyu Guo, Zhenxin Zhu, Kalok Ho, Lijun Zhou, Bohan Zeng, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wentao Zhang

机构 * Peking University(北京大学) Xiaomi EV(小米电动车) Huazhong University of Science and Technology(华中科技大学) Beijing Key Laboratory of Data Intelligence and Security (Peking University)(北京数据智能与安全重点实验室(北京大学)) Zhongguancun Academy(中关村学院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 Dream4Drive通过生成高质量的合成数据提升自动驾驶感知任务性能

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 80%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07562 2026-03-10 cs.CV 79%

Brain-WM: Brain Glioblastoma World Model

Brain-WM: 脑部胶质瘤世界模型

Chenhui Wang, Boyun Zheng, Liuxin Bao, Zhihao Peng, Peter Y. M. Woo, Hongming Shan, Yixuan Yuan

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 Brain-WM通过统一治疗预测与MRI生成,实现了肿瘤与治疗的共进化动态建模,提升了治疗计划的准确率和MRI生成的质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10480 2026-03-10 cs.CL 78%

Neuro-Symbolic Synergy for Interactive World Modeling

神经符号协同用于交互世界建模

Hongyu Zhao, Siyu Zhou, Haolin Yang, Zengyi Qin, Tianyi Zhou

专题命中 具身推理 :world model(title,abstract)

AI总结 NeSyS通过结合LLMs的概率语义先验与可执行符号规则,实现交互世界建模的表达力与鲁棒性,提升预测准确性和数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11892 2026-03-10 cs.CL 78%

R-WoM: Retrieval-augmented World Model For Computer-use Agents

R-WoM:基于检索的世界模型用于计算机使用代理

Kai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong, Dongkyu Lee, Xing Niu, Jiarong Jiang

机构 * Rutgers University(罗杰斯大学) AWS Agentic AI(亚马逊敏捷人工智能)

专题命中 具身推理 :world model(title,abstract)

AI总结 R-WoM通过整合外部检索知识提升世界模型能力,有效改善长周期模拟中的决策表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21243 2026-03-10 cs.RO 61%

RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models

RetoVLA: 通过重用寄存器令牌在视觉-语言-动作模型中实现空间推理

Jiyeon Koo, Taewan Cho, Hyunjoon Kang, Eunseom Pyo, Tae Gyun Oh, Taeryang Kim, Andrew Jaeyong Choi

机构 * School of Computing, Gachon University(高丽大学计算机学院)

专题命中 具身推理 :robotic(abstract);分类 cs.RO;robotics(journal_ref)

AI总结 RetoVLA通过重用注册令牌提升视觉-语言-动作模型的空间推理能力,实验显示在7自由度机械臂任务中平均成功率提升17.1%。

Journal ref 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07570 2026-03-10 cs.CV 57%

Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance

通过多任务自适应学习和跨维度特征引导实现高效的RGB-D场景理解

Guodong Sun, Junjie Liu, Gaoyang Zhang, Bo Wu, Yang Zhang

机构 * School of Mechanical Engineering(机械工程学院) Vehicle Measurement, Control and Safety Key Laboratory of Sichuan Province(四川省车辆测量控制与安全重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室) Shanghai Advanced Research Institute(上海先进研究院) Hubei Key Laboratory of Modern Manufacturing Quality Engineering(湖北省现代制造质量工程重点实验室)

专题命中 具身推理 :robotic(abstract);分类 cs.CV

AI总结 本文提出一种高效的RGB-D场景理解模型,通过多任务自适应学习和跨维度特征引导,提升语义分割、实例分割等任务的准确性和效率。

Comments 23 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏