arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3061 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 350 篇

2606.21775 2026-06-23 cs.LG cs.AI 新提交 81%

Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

超越下一步:用于长时程规划的变长潜世界模型

Tianqi Du, Qi Zhang, Yifei Wang, Yisen Wang

机构 * State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) Amazon AGI SF Lab(亚马逊AGI旧金山实验室) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出变长潜世界模型(VLWM),通过学习变长动作序列的条件潜状态预测,解决递归一步预测在长时程规划中的累积误差问题,结合课程训练策略,在长时程控制任务上平均提升13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21173 2026-06-23 cs.LG cs.AI 新提交 81%

Inverting the Bellman Equation: From $Q$-Values to World Models

逆推贝尔曼方程:从 $Q$ 值到世界模型

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens, Jakob N. Foerster, Oliver Richardson

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) Google DeepMind(谷歌DeepMind) Mila, University of Montreal(Mila,蒙特利尔大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文证明基于值的智能体在丰富奖励函数上训练时隐式编码世界模型,提出 $P$-learning 从 $Q$ 值提取模型,并给出编码真实转移核的充分条件,实验验证了隐式模型的准确性和泛化能力。

Comments 48 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20104 2026-06-19 cs.LG cs.AI 新提交 81%

Sensorimotor World Models: Perception for Action via Inverse Dynamics

传感器运动世界模型:通过逆动力学实现面向行动感知

Petr Ivashkov, Randall Balestriero, Bernhard Schölkopf

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Department of Computer Science, Brown University(布朗大学计算机科学系) ELLIS Institute(ELLIS研究所) ETH Zürich(苏黎世联邦理工学院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出传感器运动世界模型(SMWM),通过逆动力学正则化端到端训练潜空间世界模型,防止表示崩溃并学习与行动对齐的紧凑表示,在2D和3D控制任务中实现竞争性规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18688 2026-06-18 cs.LG cs.AI 新提交 81%

Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow

双通道接地世界建模 (DCGWM):通过异构外部接地与内向梯度流结构性防止目标干扰崩溃

Akshay Hazare

机构 * Independent Researcher(独立研究者)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出双通道接地世界建模(DCGWM),通过分区潜空间和内向梯度流,结构性防止联合嵌入预测架构中多目标接地导致的目标干扰崩溃。

Comments Position paper. Experimental validation in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18582 2026-06-18 cs.CV cs.RO eess.IV 新提交 81%

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics

ICRA 2026 GOOSE 2D细粒度语义分割挑战赛技术报告:利用DINOv3实现野外机器人中的鲁棒户外场景理解

Jaeil Park, Hyobin Choi, Sangjin Lee, Hyungtae Lim, Sung-Hoon Yoon

机构 * Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 具身推理 :robotics(title,abstract);分类 cs.RO、cs.CV

AI总结 提出一种结合DINOv3自监督骨干、ViT-Adapter和Mask2Former解码器的网络设计,以及多尺度测试增强和模型集成的推理策略,在64类细粒度越野语义分割挑战中取得第一名,复合得分76.57%。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17536 2026-06-17 cs.CV cs.AI 新提交 81%

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

OmniDrive: 一种由LLM编排的多智能体世界模型,用于多视角驾驶视频生成的统一潜在协同压缩

Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang

机构 * Peking University(北京大学) Xiamen University(厦门大学) Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) National Taiwan University(国立台湾大学) Wuhan University(武汉大学) Wuhan University of Technology(武汉理工大学) Tsinghua University(清华大学) Jimei University(集美大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出DRIVE-CHOREO,一种由LLM编排的多智能体世界模型,通过三个Qwen2.5-VL智能体协同生成位置感知的潜在序列,并利用视图-时间置换与3D VAE协同压缩,实现可控多视角视频生成,在nuScenes上达到SOTA多视角一致性和BEV mAP 21.6。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16076 2026-06-16 cs.LG cs.AI cs.GT 新提交 81%

Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting

Phys-JEPA:面向多变量时间序列预测的物理信息潜在世界模型

Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su

机构 * Tianjin University(天津大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出Phys-JEPA架构,将物理一致性约束引入潜在状态和状态转移,分解预测状态为物理和残差分量,在气候、交通、电力数据集上提升预测精度。

Comments Submitted to arXiv as a preliminary manuscript. 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15160 2026-06-16 cs.CV cs.LG 新提交 81%

DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning

DLWM: 多样化潜在世界模型用于高效多模态推理

David Huang, Lianlei Shan

机构 * University of Toronto(多伦多大学) Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出DLWM框架,结合潜在空间推理与强化学习,通过多样化潜在假设和资源感知策略提升多模态推理效率,准确率提升2-5%,内存减少24%。

Comments Preprint. 9 pages main text, 15 pages total including appendix, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14934 2026-06-16 cs.LG cs.AI 新提交 81%

Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications

可分离神经架构作为物理世界模型:从数学理论到应用

Reza T Batley, Andrew Kichline, Sourav Saha

机构 * Kevin T. Crofton Department of Aerospace and Ocean Engineering, Virginia Polytechnic Institute and State University(弗吉尼亚理工大学凯文·T·克罗夫顿航空航天与海洋工程系)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出可分离神经架构(SNA),结合神经逼近与张量分解,通过变分框架求解偏微分方程,实现高维问题代数级缩放,并在工程案例中取得显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10620 2026-06-10 cs.CV cs.AI 新提交 81%

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

图像模型能想象时间吗?ImageTime:通过时空一致性探究视觉世界建模的新基准

Xinrui Wu, Lichen Huang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出ImageTime基准,通过四关键帧协议(初始状态、动作开始、过渡状态、最终状态)评估图像生成模型在时空一致性上的表现,揭示模型在维持连贯视觉世界状态方面的能力与不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09803 2026-06-09 cs.CV cs.GR cs.LG 新提交 81%

Echo-Memory: A Controlled Study of Memory in Action World Models

Echo-Memory:动作世界模型中记忆的受控研究

Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan

机构 * Joy Future Academy(京东探索研究院)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出Echo-Memory框架,通过控制变量法研究动作条件世界模型中的记忆机制,发现原始上下文容量和块状状态空间递归对开放域返回任务至关重要。

Comments 9 figures and 28 pages, Code at \href{https://github.com/Echo-Team-Joy-Future-Academy-JD/Echo-Memory}{this URL}

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07974 2026-06-09 cs.RO cs.AI 新提交 81%

PRISM: PRior-guided Imagination Sampling in world Models

PRISM:世界模型中基于先验引导的想象采样

Yuhai Wang, Jiawei Xia, Rongxuan Zhou, Xiao Hu, Yongliang Shi, Jing Du, Yang Ye

机构 * Northeastern University(东北大学) University of California, Berkeley(加州大学伯克利分校) Qiyuan Lab(启元实验室) University of Florida(佛罗里达大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出PRISM框架,通过从世界模型编码器提取状态条件高斯先验,并利用精度加权高斯乘积更新规划器的采样分布,在不增加架构复杂度的情况下显著提升基于模型的连续控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10362 2026-07-14 cs.LG 新提交 80%

A Control Theory of Predictability in Latent World Models

潜在世界模型中可预测性的控制理论

Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo

机构 * University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology(香港科技大学) Hong Kong Baptist University(香港浸会大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 研究潜在世界模型中可预测性的控制理论,指出当前以预测误差为目标不可靠,重新定义目标为预测与真实计划成本的差异,证明规划器次优性受其限制,通过实验证实相关结论。

Comments Preprint about latent world models, Koopman operator and control theory, 33 pages, 1 figure. Main text about 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22729 2026-06-23 cs.RO 新提交 80%

Temporal Logic Guidance for Action-Only Diffusion Policies with World Models

基于时序逻辑的动作扩散策略与世界模型引导

Moritz Zoellner, Anastasios Manganaris, Rohan Paleja

机构 * German Academic Exchange Service (DAAD)(德国学术交流中心(DAAD))

专题命中 具身推理 :world model(title,abstract);分类 cs.RO;robot learning(comments)

AI总结 提出一种利用世界模型实现时序逻辑鲁棒性可微评估的引导方法,在不重新训练的情况下改善动作扩散策略的约束满足,在Robomimic任务中将违规率从80%降至4%。

Comments Accepted at the ICRA 2026 Workshop on Bridging the Gap between Robot Learning and Human-Robot Interaction. 3 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19383 2026-06-19 cs.RO cs.CV 新提交 80%

3D Scene Graphs: Open Challenges and Future Directions

3D场景图:开放挑战与未来方向

Dennis Rotondi, Francesco Argenziano, Sebastian Koch, Nathan Hughes, Martin Buechner, Johanna Wald, Lukas Rosenberger Schmid, Daniele Nardi, Abhinav Valada, Liam Paull, Federico Tombari, Luca Carlone, Kai O. Arras

机构 * University of Stuttgart(斯图加特大学) IMPRS-IS(马克斯·普朗克研究所-智能系统) Sapienza University of Rome(罗马萨皮恩扎大学) Google(谷歌) MIT(麻省理工学院) University of Freiburg(弗赖堡大学) UTN University of Montreal(蒙特利尔大学UTN分校) Mila TU Munich(慕尼黑技术大学Mila)

专题命中 具身推理 :robotics(abstract,comments);manipulation(abstract);navigation(abstract);分类 cs.RO、cs.CV

AI总结 本文统一综述3D场景图(3DSG)的构建、应用与评估,分析现有建模选择与开放挑战,旨在推动鲁棒部署。

Comments Invited article for the Annual Review of Control, Robotics, and Autonomous Systems Volume 10

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18077 2026-08-19 cs.RO 新提交 79%

Hydra-0: Action Flow for Generalist World Modeling and Control

Hydra-0:面向通用世界建模与控制的动作流

Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang

机构 * NVIDIA(英伟达) Brown University(布朗大学) Columbia University(哥伦比亚大学) Harvard University(哈佛大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO

AI总结 本研究提出Hydra-0通用世界模型,以动作流为条件实现跨实体、任务等的通用建模控制,在RoboLab基准获高相关性,还涌现出逆模式,可助力机器人控制。

Comments Project page: this https URL (https://nvidia-isaac.github.io/video_to_data/hydra-0/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16955 2026-08-19 cs.MA cs.LG 新提交 79%

WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization

WONDER:一种基于无线电世界模型的多智能体无人机覆盖优化协商框架

Jiahao Huang, Rongpeng Li, Zhifeng Zhao, Guoru Ding, Honggang Zhang

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 针对灾后无线覆盖中断问题,提出基于JEPA无线电世界模型与PPO演员交替更新的WONDER框架,在含62个城市场景的RadioDynamics仿真中,其平衡得分0.870,覆盖优于STACCA且保持无人机全连通。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15043 2026-08-18 cs.AI 新提交 79%

SCOPE: Score-Isolated Agentic Optimization for Video World Models

SCOPE:用于视频世界模型的分数隔离智能体优化

Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao

机构 * Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Tencent(腾讯)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本研究针对视频世界模型推理时改进的评估难题,提出SCOPE框架,通过类型化状态更新与冻结策略实现可审计适应,在Physics-IQ基准上较冻结基线提升+14.24,同时揭示推理时更新的益处需原则性部署机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12939 2026-08-14 cs.LG 新提交 79%

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

用动作条件预测一致性诊断JEPA世界模型

Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian

机构 * Huawei(华为) University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 本研究针对JEPAs世界模型易受视觉扰动影响的问题,提出动作条件预测一致性(ACPC)诊断方法,定义IR与SR指标,经四视觉控制任务实验验证其可预测扰动带来的预测及代价变化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12564 2026-08-14 cs.LG 新提交 79%

Scaling Automatic Research Agents via World Models

通过世界模型扩展自动研究智能体

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Chenlei Guo, Jingrui He, Zhenyu Liao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 针对自动研究智能体扩展时的训练瓶颈,提出WMRL方法,结合两种缓解措施提升收敛性,训练加速3-4倍且性能优于更大规模智能体,还可迁移至多类任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11601 2026-08-13 cs.CV 新提交 79%

How Can Driving World Models Do Counterfactual Prediction?

驾驶世界模型如何能进行反事实预测?

Jiaru Zhang, Can Cui, Yi Xu, Xin Ye, Ruqi Zhang, Ziran Wang

机构 * Purdue University(普渡大学) Bosch Center for Artificial Intelligence(博世人工智能中心)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文指出驾驶世界模型的反事实预测目标与直接动作条件预测存在根本差距,构建基准验证该问题,并提出简单无训练流程提升反事实预测效果,呼吁开发更好的相关方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10107 2026-08-12 cs.CV 新提交 79%

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

4D-WAM:面向自动驾驶的4D一致世界建模

Jiacheng Fu, Yibo Yuan, Meng Tian, Yue Li, Jiangtong Zhu, Jianhua Han, Yueyi Zhang, Jianwu Fang, Jianru Xue, Hang Xu, Zhiwei Xiong

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出4D-WAM模型,通过几何基础模型的训练时监督与面向决策的时间步长采样策略,提升自动驾驶中世界-动作模型的4D场景一致性,在NAVSIM-v1、NAVSIM-v2基准上实现最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09926 2026-08-11 cs.CV 新提交 79%

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

学习世界如何演化:基于潜动态推理的外推视频世界模型

Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji, Jiahan Zhang, Xin Lin, Haolin Lu, Zhe Lin, Manmohan Chandraker

机构 * UCSD(加利福尼亚大学圣迭戈分校) Adobe(奥多比公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 针对主流视频扩散模型未建模像素时间转换的问题,提出LDR方法,在PhyWorld基准上实现更优动态外推,参数更少、速度更快,是首个能泛化至训练分布外的视频世界模型

Comments Project page: https://lat-dyn-reason.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09537 2026-08-11 cs.AI 新提交 79%

verdi: retrieval is not transfer for continual world model optimization

VERDI:检索并非持续世界模型优化的迁移

Junyu Wu, Shiqin Nie, Youyi Kou, Baohua Yin, Guocai Yao, Qingyu Chen, Jingheng Ma, Shiji Zhou, Hongyong Song, Mingchen Zhuge, Sen Cui, Changshui Zhang

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文针对基础世界模型优化中策略难迁移的问题,提出VERDI框架,通过构建优化指纹、检索假设并经目标侧验证来实现持续优化,可降低搜索与GPU成本,减少负迁移并提升迁移结果预测准确率。

Comments 28pages, 13figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08982 2026-08-11 cs.LG 新提交 79%

Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models

双回滚:交互式视频世界模型中的噪声耦合反事实分支

Yu Ma, Hongli Shi, Xinran Xu

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 本研究提出噪声耦合双回滚框架,解决交互式视频世界模型的反事实生成问题,规避近似逆问题,定义时空局部性度量,待开展实验验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08689 2026-08-11 cs.AI 新提交 79%

A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration

结构动力学图世界模型:统一建模、约束展开与可解释校准

Wei Wang, Yaosen Chen, Han Yang, Yuegen Liu, Mingli Luo, Xinxin Jiao, Xuming Wen, Ming Liu

机构 * Sobey(Sobey公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 提出SD-GWM结构动力学图世界模型,实现异构集成、语义保真与可审计治理,在极端洪水预测中较基线模型有8-28倍精度提升,为可审计的时空挖掘提供约束安全的可验证底物。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07981 2026-08-11 cs.CV 新提交 79%

Distilling Physical Priors into Streaming World Models

将物理先验知识蒸馏为流式世界模型

Liangliang Zhao, Junying Wang, Danni Yang, Yifan Chang, Bin Fu, Yu Qiao, Bowen Zhou, Yihao Liu

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出PhyS三阶段框架,构建120K物理交互数据集,经微调、蒸馏及在线强化学习结合TCR方法,提升流式世界模型的物理一致性,在PhysicsIQ等基准上取得显著性能提升。

Comments 9 pages, 7 figures. Project page: https://lyongo.github.io/PhyS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07809 2026-08-11 cs.AI 新提交 79%

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

CausalNav:用于物理参数偏移下控制的可靠性可验证因果世界模型

Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen

机构 * University of Wollongong(卧龙岗大学) CSIRO’s Data61(联邦科学与工业研究组织Data61) Adelaide University(阿德莱德大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出CausalNav控制器,其结合因果转移图与多门控可靠性验证机制,在物理参数偏移的CartPole-v1和Pendulum-v1任务上,较9种基准取得最佳平均排名,关键在于经认证的弃权(不执行)提升了部署安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07409 2026-08-10 cs.CV 新提交 79%

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

UniJEPA:面向任务无关视觉世界建模的统一联合嵌入预测架构

An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 UniJEPA是统一JEPAs,共享潜在空间联合学习图像级与视频级预测,抗坍塌,经后训练可零样本规划,在多基准上性能相当或更优,规划速度更快。

Comments 10 pages, 7 figures; Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07107 2026-08-10 cs.AI 新提交 79%

MemWM: Memory-Augmented Text-Based World Model

MemWM:记忆增强的基于文本的世界模型

Yujun Wang, Tao Zhang, Jinhe Bi, Aniri, Wenxuan Ye, Boliang Liu, Sikuan Yan, Shuning Wang, Xuebing Zhou, Sören Pirk, Hinrich Schütze, Yunpu Ma

机构 * LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Huawei Heisenberg Research Center(华为海森堡研究中心) Zhejiang University(浙江大学) Technical University of Munich (TUM)(慕尼黑工业大学) TU Berlin(柏林工业大学) Kiel University(基尔大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 该研究提出记忆增强的基于文本的世界模型MemWM,通过引入世界记忆解决世界模型的系统性预测错误,在ALFWorld等基准上提升智能体规划成功率与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏