arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3091 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3091 篇

2512.13821 2026-02-06 cs.LG 74%

The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces

代码世界模型的双重生命:通过执行轨迹证明性地揭示恶意行为

Subramanyam Sahoo

机构 * Berkeley AI Safety Initiative (BASIS) UC Berkeley(伯克利人工智能安全倡议(BASIS)加州大学伯克利分校)

专题命中 具身推理 :world model(title);分类 cs.LG

AI总结 通过语义轨道分析,CTVP验证代码生成模型的恶意行为,引入ARQ量化验证成本,理论证明对抗鲁棒性非游戏化。

Comments 13 Pages, A Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03146 2026-02-04 cs.AI 74%

General Agents Contain World Models, even under Partial Observability and Stochasticity

通用智能体包含世界模型,即使在部分可观测性和随机性下

Santiago Cifuentes

机构 * Dovetail Research(Dovetail研究)

专题命中 具身推理 :world model(title);分类 cs.AI

AI总结 本研究证明了即使在部分可观测性和随机性下,通用智能体仍需学习环境模型。

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14514 2026-01-22 cs.AI q-bio.NC 74%

"Just in Time" World Modeling Supports Human Planning and Reasoning

即时世界建模支持人类规划与推理

Tony Chen, Sam Cheyette, Kelsey Allen, Joshua Tenenbaum, Kevin Smith

机构 * MIT Department of Brain and Cognitive Sciences(麻省理工学院脑科学与认知科学系) UBC Departments of Computer Science and Psychology(不列颠哥伦比亚大学计算机科学与心理学系)

专题命中 具身推理 :world model(title);分类 cs.AI

AI总结 本文提出'即时'框架,通过在线构建简化表示实现高效心理模拟,支持人类规划与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14354 2026-01-22 cs.LG 74%

VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models

VJEPA:变分联合嵌入预测架构作为概率世界模型

Yongchao Huang

专题命中 具身推理 :world model(title);分类 cs.LG

AI总结 VJEPA通过变分目标学习预测分布,统一表征学习与贝叶斯过滤,提供概率世界模型框架,适用于高维噪声环境中的鲁棒规划。

Comments 77 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08536 2025-11-12 cs.CV 74%

3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation

Yunhong He, Zhengqing Yuan, Zhengzhong Tu, Yanfang Ye, Lichao Sun

专题命中 具身推理 :world model(title);分类 cs.CV

Comments Accepted by AAAI 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04363 2025-05-16 cs.AI 74%

AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents

Petr Anokhin, Nikita Semenov, Artyom Sorokin, Dmitry Evseev, Andrey Kravchenko, Mikhail Burtsev, Evgeny Burnaev

专题命中 具身推理 :world model(title);分类 cs.AI

Comments Code for this work is avaliable at https://github.com/AIRI-Institute/AriGraph

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07375 2025-02-26 cs.CV 74%

StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization

Jinlu Zhang, Jiji Tang, Rongsheng Zhang, Tangjie Lv, Xiaoshuai Sun

专题命中 具身推理 :world model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08268 2025-02-05 cs.LG 74%

World Model on Million-Length Video And Language With Blockwise RingAttention

Hao Liu, Wilson Yan, Matei Zaharia, Pieter Abbeel

专题命中 具身推理 :world model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12671 2024-11-20 cs.AI cs.CL cs.ET 74%

Neurosymbolic Graph Enrichment for Grounded World Models

Stefano De Giorgis, Aldo Gangemi, Alessandro Russo

专题命中 具身推理 :world model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10875 2024-03-19 cs.LG 74%

Probabilistic World Modeling with Asymmetric Distance Measure

Meng Song

专题命中 具身推理 :world model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16972 2024-01-31 cs.CV eess.IV 74%

Deep 3D World Models for Multi-Image Super-Resolution Beyond Optical Flow

Luca Savant Aira, Diego Valsesia, Andrea Bordone Molini, Giulia Fracastoro, Enrico Magli, Andrea Mirabile

专题命中 具身推理 :world model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14159 2024-01-26 cs.CV 74%

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, Lei Zhang

专题命中 具身推理 :world model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14909 2023-11-03 cs.AI 74%

Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning

Lin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao Kambhampati

专题命中 具身推理 :world model(title);分类 cs.AI

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13394 2023-10-23 cs.CL cs.AI cs.CY 74%

POSQA: Probe the World Models of LLMs with Size Comparisons

Chang Shu, Jiuzhou Han, Fangyu Liu, Ehsan Shareghi, Nigel Collier

专题命中 具身推理 :world model(title);分类 cs.AI

Comments Accepted by EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05746 2023-01-18 cs.CL cs.AI 74%

Infusing Commonsense World Models with Graph Knowledge

Alexander Gurung, Mojtaba Komeili, Arthur Szlam, Jason Weston, Jack Urbanek

专题命中 具身推理 :world model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11371 2022-07-04 cs.LG 74%

Learning Symmetric Embeddings for Equivariant World Models

Jung Yeon Park, Ondrej Biza, Linfeng Zhao, Jan Willem van de Meent, Robin Walters

专题命中 具身推理 :world model(title);分类 cs.LG

Comments ICML 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04585 2022-06-22 cs.RO cs.CL 74%

Extracting Zero-shot Common Sense from Large Language Models for Robot 3D Scene Understanding

William Chen, Siyi Hu, Rajat Talak, Luca Carlone

专题命中 具身推理 :robotics(abstract,comments);robotic(abstract);分类 cs.RO;robot learning(comments)

Comments 4 pages (excluding references and appendix), 2 figures, 2 tables. Submitted to Robotics: Science and Systems 2022 2nd Workshop on Scaling Robot Learning. Corrected typos and notation

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09487 2020-03-31 cs.CV 74%

A Robotic 3D Perception System for Operating Room Environment Awareness

Zhaoshuo Li, Amirreza Shaban, Jean-Gabriel Simard, Dinesh Rabindran, Simon DiMaio, Omid Mohareri

专题命中 具身推理 :robotic(title);分类 cs.CV

Comments Accepted in IPCAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.08040 2019-07-19 cs.LG cs.NE stat.ML 74%

Convolutional Reservoir Computing for World Models

Hanten Chang, Katsuya Futagami

专题命中 具身推理 :world model(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05215 2026-08-07 cs.RO cs.CV 新提交 73%

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

VLAff:用于统一可操作 affordance 的视觉-语言-affordance 模型

Jihoon Oh, Kento Kawaharazuka, Kei Okada

机构 * University of Tokyo(东京大学)

专题命中 具身推理 :robot learning(abstract);manipulation(abstract);分类 cs.RO、cs.CV

AI总结 本研究提出 VLAff 模型,结合 EgoAffordance 数据集,解决人类与机器人 embodiment 不匹配问题,实现视觉 affordance 预测及真实机器人零样本操作等应用。

Comments 8 pages, 5 figures. Accepted to IEEE/RSJ IROS 2026. Project page: https://ojh6404.github.io/vlaff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27036 2026-07-30 cs.CV cs.LG 新提交 73%

Mitigating Compounding Error via Video Representation Regularization

通过视频表示正则化缓解复合误差

Taiye Chen, Qi Zhang, Yisen Wang

机构 * Peking University(北京大学)

专题命中 具身推理 :robotics(abstract);world model(abstract);分类 cs.CV、cs.LG

AI总结 针对视频扩散世界模型自回归生成的复合误差问题,研究发现其与表示维度崩溃相关,提出视频表示正则化方法,在VBench指标上显著优于Diffusion Forcing,提升了长视频生成的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26903 2026-07-30 cs.AI cs.RO 新提交 73%

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

从被动视频到可编辑体验:具身智能的物理基础体验合成

Jia Luo

专题命中 具身推理 :embodied AI(abstract);manipulation(abstract);分类 cs.RO、cs.AI

AI总结 针对具身AI的数据瓶颈,提出Pegasus低资源框架,通过结构化知识传递将人类操作视频转化为机器人可学习数据,经多基准与机器人评估验证其跨具身翻译及数据生成的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11689 2026-07-14 cs.RO cs.AI 新提交 73%

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

从世界行动模型到具身大脑:开放世界物理智能路线图

Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, Xuelong Li

机构 * IEEE

专题命中 具身推理 :embodied agent(abstract);world model(abstract);分类 cs.RO、cs.AI

AI总结 研究针对通用人工智能中物理世界推理行动的进展零散问题,提出以具身大脑为中心的物理智能协同进化路线图,利用世界行动模型等,通过物理框架、共享契约和闭环训练等构建模块化智能栈,推动物理智能发展。

Comments Ongoing work

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04419 2026-07-14 cs.CL cs.AI cs.LG 版本更新 73%

Context-Dependent Affordance Computation in Vision-Language Models

视觉-语言模型中基于情境的可及性计算

Murad Farzulla

机构 * Dissensus AI King’s College London(伦敦国王学院)

专题命中 具身推理 :robotics(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 本研究揭示视觉-语言模型中可及性计算的高度情境依赖性,通过大规模实验显示超过90%的词汇描述受情境影响,提出动态本体投影方法替代静态世界建模。

Comments 33 pages, 13 tables, 3 figures. Code available at: https://github.com/studiofarzulla/semantic-vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02542 2026-07-07 cs.AI cs.CV 新提交 73%

iFLYTEK-Embodied-Omni Technical Report

科大讯飞-具身全能技术报告

Yuan Zhang, Jingfei Ni, Guanchen Lu, Shiqi Zhang, Qingshan Xu, Chi Liu, Xin Nie, Wenjie Xu, Lin Gao, Zhiyuan Cheng, Mingxin Zhou, Jiajia Wu, Diyuan Liu, Jia Pan, Chao Ji

机构 * iFLYTEK LindenBot University of Science and Technology of China(中国科学技术大学)

专题命中 具身推理 :embodied agent(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 研究通用具身智能体,提出统一多模态基础模型iFLYTEK-Embodied-Omni,通过共享多模态自注意力联合建模视觉、语言和动作,结合多种数据构建数据集并采用四阶段策略训练,实现脑-小脑协作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26964 2026-06-29 cs.AI cs.CV 新提交 73%

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Look-Before-Move:动态3D故事世界中的叙事驱动世界视觉注意力

Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, Huadong Mo, Zhenhong Sun

专题命中 具身推理 :embodied AI(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 提出Look-Before-Move框架,通过语义观察合约、蒙特卡洛视点搜索和语义轨迹接地,在动态3D故事世界中实现叙事驱动的视觉注意力规划,提升主体感知、意图一致性和轨迹质量。

Comments 25 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20905 2026-06-23 cs.RO cs.AI 新提交 73%

Vesta: A Generalist Embodied Reasoning Model

Vesta: 一种通用具身推理模型

Johan Bjorck, Zhiqi Li, Yunze Man, Jing Wang, An-Chieh Cheng, Sifei Liu, Shihao Wang, Zhiding Yu, Abhishek Badki, Stan Birchfield, Valts Blukis, Yevgen Chebotar, Siyi Chen, Sicong Leng, Yu-Cheng Chou, Tianli Ding, Boyi Li, Zhengyi Luo, Hang Su, Jonathan Tremblay, Tingwu Wang, Bowen Wen, Jimmy Wu, Xianghui Xie, Hanrong Ye, Hongxu Yin, K. R. Zentner, Liangyan Gui, Yu-Xiong Wang, Yuke Zhu, Linxi "Jim" Fan, Jan Kautz

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, San Diego(加州大学圣地亚哥分校) The Hong Kong Polytechnic University(香港理工大学) University of Michigan(密歇根大学) Nanyang Technological University(南洋理工大学) Johns Hopkins University(约翰霍普金斯大学) University of Tübingen(图宾根大学)

专题命中 具身推理 :navigation(abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 提出统一的基础模型Vesta,整合定位、空间推理、导航和长时规划能力,通过大规模空间感知语料库和简单多模态记忆机制,在多个基准上平均优于专用模型20%以上,在真实机器人任务中成功率提升35%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12688 2026-06-16 cs.LG cs.AI cs.DC 新提交 73%

M*: A Modular, Extensible, Serving System for Multimodal Models

M*: 一个模块化、可扩展的多模态模型服务系统

Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身推理 :world model(abstract);robotic(abstract);分类 cs.AI、cs.LG

AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。

Comments The codebase is available at https://github.com/mstar-project/mstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00104 2026-06-02 cs.RO cs.AI 73%

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

PEACE: 一种用于无人机的带约束执行的规划-执行智能体

Erdem Uysal, Timo Kehrer, Sebastiano Panichella

机构 * Institute of Computer Science, University of Bern(伯尔尼大学计算机科学研究所) AI4I - The Italian Institute of Artificial Intelligence(意大利人工智能研究所)

专题命中 具身推理 :robotics(abstract);world model(abstract);分类 cs.RO、cs.AI

AI总结 提出一种基于大语言模型的规划-执行智能体架构,通过解耦高层任务规划与低层控制,并引入约束执行层和有限重规划,实现无人机可解释、可约束的自主飞行。

Comments Accepted to ICRA 2026 Workshop on Semantics for Reliable Robot Autonomy: From Environment Understanding and Reasoning to Safe Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13169 2026-05-18 cs.CV cs.AI 73%

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World

PanoWorld:迈向360度全景世界的空间超感知

Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu, Zhen Wang, Donglian Qi, Yunfeng Yan, Xi Chen

机构 * Zhejiang University(浙江大学) University of California, San Diego(加州大学圣地亚哥分校) University of California, Irvine(加州大学伊维特分校) The University of Hong Kong(香港大学)

专题命中 具身推理 :navigation(abstract);robotic(abstract);分类 cs.AI、cs.CV

AI总结 本文提出PanoWorld,通过构建全景原生理解能力,解决传统多模态大模型在空间感知上的不足,通过全景空间交叉注意力机制提升3D空间推理能力,并建立PanoSpace-Bench基准测试,验证了全景原生监督的有效性。

Comments Project page: https://wcpcp.github.io/PanoWorld

详情

展开后加载摘要…

URL PDF HTML 收藏