arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 504 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 具身与机器人 504 篇

2602.10983 2026-02-13 cs.RO 88%

Scaling World Model for Hierarchical Manipulation Policies

为分层操作策略扩展世界模型

Qian Long, Yueze Wang, Jiaxi Song, Junbo Zhang, Peiyan Li, Wenxuan Wang, Yuqi Wang, Haoyang Li, Shaoxuan Xie, Guocai Yao, Hanbo Zhang, Xinlong Wang, Zhongyuan Wang, Xuguang Lan, Huaping Liu, Xinghang Li

机构 * Xi’an Jiao Tong University(西安交通大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 本文提出分层VLA框架,利用大规模预训练世界模型提升机器人操作在分布外场景的泛化能力,实验显示性能提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06990 2025-08-12 cs.RO 88%

Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

Yue Hu, Junzhe Wu, Ruihan Xu, Hang Liu, Avery Xi, Henry X. Liu, Ram Vasudevan, Maani Ghaffari

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17491 2025-07-24 cs.RO 88%

X-MOBILITY: End-To-End Generalizable Navigation via World Modeling

Wei Liu, Huihua Zhao, Chenran Li, Joydeep Biswas, Billy Okal, Pulkit Goyal, Yan Chang, Soha Pouya

机构 * NVIDIA UC Berkeley(加州大学伯克利分校) UT Austin(德克萨斯大学奥斯汀分校)

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08750 2025-07-15 cs.RO 88%

DexSim2Real$^{2}$: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation

Taoran Jiang, Yixuan Guan, Liqian Ma, Jing Xu, Jiaojiao Meng, Weihang Chen, Zecui Zeng, Lusong Li, Dan Wu, Rui Chen

机构 * Department of Mechanical Engineering, Tsinghua University(清华大学机械工程系) Institute for Robotics and Intelligent Machines (IRIM), Georgia Institute of Technology(机器人与智能机器研究所(IRIM),佐治亚理工学院) JD Explore Academy, Beijing, China(京东探索研究院,北京,中国)

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

Comments Project Webpage: https://jiangtaoran.github.io/dexsim2real2web/ . arXiv admin note: text overlap with arXiv:2302.10693

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19957 2026-05-20 cs.CV cs.AI cs.RO 87%

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

为混合具身体验中的长时域演化构建世界-自我模型

Zuyao Lin, Jianhui Zhang, Peidong Jia, Xiaoguang Zhao, Shanghang Zhang, Xingyu Chen

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);embodied world model(abstract);world model(abstract)

AI总结 本文提出了一种新的世界-自我建模范式,通过分解未来演化为世界和自我组件,解决混合任务中长时域具身体验中的退化问题,并通过HTEWorld基准测试验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08546 2026-03-10 cs.RO cs.CV cs.LG 87%

Interactive World Simulator for Robot Policy Training and Evaluation

交互式世界模拟器用于机器人策略训练与评估

Yixuan Wang, Rhythm Syed, Fangyu Wu, Mengchao Zhang, Aykut Onol, Jose Barreiros, Hooshang Nayyeri, Tony Dear, Huan Zhang, Yunzhu Li

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 交互式世界模拟器通过构建物理一致的世界模型,实现了高效的机器人策略训练与评估,支持长时间稳定交互并展现与现实世界性能的强相关性。

Comments Project Page: https://yixuanwang.me/interactive_world_sim

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11643 2026-07-14 cs.RO cs.AI 新提交 87%

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

小米机器人-U0:基于世界基础模型的统一具身合成

Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li

机构 * Xiaomi Robotics(小米机器人)

专题命中 具身与机器人 :world model(abstract);world models(abstract);embodied world model(abstract);world model(abstract)

AI总结 研究针对基础模型应用于具身场景受限问题,提出小米机器人-U0多模态自回归模型,将具身生成扩展为基础图像和视频生成的形式,联合优化多种生成任务,在多方面取得领先成果,证明基础世界模型可用于具身智能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18610 2026-06-29 cs.RO cs.CV 新提交 87%

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

SC3-Eval: 通过自洽视频生成评估机器人基础模型

Wei-Cheng Tseng, Gashon Hussein, Yuzhu Dong, Allen Z. Ren, Lucy X. Shi, XuDong Wang, Sergey Levine, Zhaoshuo Li, Jinwei Gu, Florian Shkurti, Ming-Yu Liu, Quan Vuong

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) NVIDIA(英伟达) Physical Intelligence Stanford University(斯坦福大学) UC Berkeley(加州大学伯克利分校) Allen Institute for AI(艾伦人工智能研究所)

专题命中 具身与机器人 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出SC3-Eval方法,利用前向-反向动力学一致性、跨视角一致性和测试时一致性,将预训练视频基础模型转化为准确的策略评估器,在7个真实世界策略上达到0.929的皮尔逊相关系数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13494 2026-06-12 cs.RO cs.CV 新提交 87%

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

NavWAM:用于目标条件视觉导航的导航世界动作模型

Daichi Azuma, Taiki Miyanishi, Koya Sakamoto, Shuhei Kurita, Yaonan Zhu, Petr Khrapchenkov, Motoaki Kawanabe, Yusuke Iwasawa, Yutaka Matsuo

机构 * The University of Tokyo(东京大学) National Institute of Informatics(国立信息学研究所) AIRoA ATR

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出NavWAM,一种扩散变换器策略,通过联合学习未来观测、目标进度值和动作块,将导航世界模型预测直接转化为可执行动作,在离线基准和真实机器人部署中优于基于规划的世界模型基线。

Comments Project page: https://dachii-azm.github.io/navwam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05960 2026-06-05 cs.RO 86%

Towards a Data Flywheel for Embodied Intelligence in Logistics

面向物流具身智能的数据飞轮

Anlan Yu, Zaishu Chen, Zhiqing Hong, Daqing Zhang

机构 * Peking University(北京大学) JD Logistics(京东物流) HKUST (Guangzhou)(香港科技大学(广州))

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出一种数据驱动的物流具身智能框架,通过构建数据飞轮将日常操作转化为可复用数据资产,利用世界模型生成长尾包裹操作的可靠监督,并整合多模态数据实现策略持续改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21797 2026-02-11 cs.CV 86%

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

MoWM: 通过潜在到像素特征调制的混合世界模型实现具身规划

Yangcheng Yu, Xin Jin, Yu Shang, Xin Zhang, Haisheng Su, Wei Wu, Yong Li

机构 * Tsinghua University(清华大学) Manifold AI Shanghai Jiao Tong University(上海交通大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 MoWM通过融合潜在世界模型与像素特征,提升具身规划中动作解码的精度与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07823 2026-01-13 eess.SY cs.RO cs.SY 86%

Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions

机器人中的视频生成模型——应用、研究挑战与未来方向

Zhiting Mei, Tenny Yin, Ola Shorinwa, Apurva Badithela, Zhonghe Zheng, Joseph Bruno, Madison Bland, Lihan Zha, Asher Hancock, Jaime Fernández Fisac, Philip Dames, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学) Temple University(Temple 大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);embodied world model(abstract);world model(abstract)

AI总结 本文综述了视频生成模型在机器人学中的应用、研究挑战及未来方向,探讨了其在物理模拟、动作预测和策略评估中的作用及面临的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18647 2026-08-20 cs.RO cs.LG 新提交 85%

Progressive Experience Fusion for Multi-Task World Model Control in Endovascular Navigation

用于血管内导航多任务世界模型控制的渐进式经验融合

Harry Robertshaw, Maxence Boels, Nikola Fischer, Sebastien Ourselin, Christos Bergeles, Alejandro Granados, Thomas C Booth

机构 * School of Biomedical Engineering & Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院) Surgical & Interventional Engineering(外科与介入工程)

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.LG、cs.RO

AI总结 本研究提出渐进式经验融合(PEF)训练多任务TD-MPC2控制器,在血管内导航任务中提升了成功率,可实现跨血管结构迁移与患者特异性微调,为临床应用提供概念验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26657 2026-08-07 cs.RO cs.CV 版本更新 85%

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

Enfold:将世界生成器计算融入预测表示以实现高效具身控制

Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan

机构 * Shanghai Jiao Tong University(上海交通大学) South China University of Technology(华南理工大学) Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 Enfold将世界生成器计算融入预测表示,在LIBERO等多任务中实现高效具身控制,大幅降低动作延迟,且能适应场景变化,为具身控制提供了新范式。

Comments project page, https://zwl666666.github.io/enfold/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07392 2026-07-09 cs.LG cs.IR cs.RO 版本更新 85%

Event-Centric World Modeling with Memory-Augmented Retrieval for Embodied Decision-Making

基于记忆增强检索的事件中心世界建模用于具身决策

Zhaowen Fan, Rongchao Zhang, Yunxiang Han

机构 * Zhaowen Fan(Fan 研究院) Rongchao Zhang(Zhang 实验室)

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.LG、cs.RO

AI总结 本文提出一种事件中心世界建模框架,通过记忆增强检索实现具身决策,结合语义事件和先前经验进行决策,提升动态环境的结构抽象与可解释性。

Comments 12 pages, 9 figures, 4 tables, currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14135 2026-05-25 cs.RO cs.CV 85%

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation

GAF: 高斯动作场作为机器人操作中动态世界建模的4D表示

Ying Chai, Litao Deng, Ruizhi Shao, Jiajun Zhang, Kangchen Lv, Liangjun Xing, Xiang Li, Hongwen Zhang, Yebin Liu

机构 * Tsinghua University(清华大学) Beijing Normal University(北京师范大学) Shadow AI

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 提出GAF,通过扩展3D高斯溅射引入可学习运动属性,实现动态场景的4D建模,并利用动作-视觉对齐去噪框架提升机器人操作成功率。

Comments https://ChaiYing1.github.io/projects/GAF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16669 2026-03-18 cs.RO cs.CV 85%

Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation

Kinema4D: 用于空间时间具身模拟的运动学4D世界建模

Mutian Xu, Tianbao Zhang, Tianqi Liu, Zhaoxi Chen, Xiaoguang Han, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S-Lab) SSE, CUHKSZ(香港中文大学深圳校区SSE)

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 本文提出Kinema4D,一种基于动作条件的4D生成机器人模拟器,通过精确的4D机器人控制表示和环境反应的生成建模,实现了高精度的空间时间交互模拟,并首次展示零样本迁移能力。

Comments Project page: https://mutianxu.github.io/Kinema4D-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11797 2026-07-07 cs.RO cs.CV 版本更新 84%

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

AnchorDream:将视频扩散用于具身感知机器人数据合成

Junjie Ye, Rong Xue, Basile Van Hoorick, Pavel Tokmakov, Muhammad Zubair Irshad, Yue Wang, Vitor Guizilini

机构 * Toyota Research Institute(丰田研究院) USC Physical Superintelligence (PSI) Lab(USC物理超智能(PSI)实验室)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对机器人示范数据收集瓶颈,AnchorDream利用预训练视频扩散模型,通过机器人运动渲染来合成数据,无需环境建模,从少量示范生成多样高质量数据集,提升下游策略学习。

Comments Project page: https://jay-ye.github.io/AnchorDream/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27817 2026-05-28 cs.RO cs.AI cs.CV cs.LG 84%

Turning Video Models into Generalist Robot Policies

将视频模型转化为通用机器人策略

Sizhe Lester Li, Evan Kim, Xingjian Bai, Tong Zhao, Tao Pang, Max Simchowitz, Vincent Sitzmann

机构 * MIT(麻省理工学院) CMU(卡内基梅隆大学) Amazon FAR(亚马逊公司)

专题命中 具身与机器人 :world model(abstract);video world model(abstract);world model(abstract);video world model(abstract)

AI总结 提出一种解耦的视频到动作策略VERA,利用无动作视频世界模型和基于机器人雅可比矩阵的逆动力学模型,实现跨本体的零样本机器人控制。

Comments project page: https://vera.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06886 2025-08-26 cs.CV cs.AI cs.LG cs.MA cs.RO 84%

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, Liang Lin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Sun Yat-sen University(中山大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) Peng Cheng Laboratory(鹏城实验室) Institute of Digital Media, Peking University(北京大学数字媒体研究院)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments The comprehensive review of Embodied AI. We also provide the resource repository for Embodied AI: https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10788 2024-06-18 cs.RO 84%

Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics

Jad Abou-Chakra, Krishan Rana, Feras Dayoub, Niko Sünderhauf

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.10138 2018-03-01 cs.RO 84%

Real-World Modeling of a Pathfinding Robot Using Robot Operating System (ROS)

Sayyed Jaffar Ali Raza, Nitish A. Gupta, Nisarg Chitaliya, Gita R. Sukthankar

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15109 2024-12-20 cs.RO 84%

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Yang Tian, Sizhe Yang, Jia Zeng, Ping Wang, Dahua Lin, Hao Dong, Jiangmiao Pang

专题命中 具身与机器人 :world model(abstract);world models(abstract);dynamics model(title,abstract);world model(abstract)

Comments Project page: https://nimolty.github.io/Seer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23181 2026-07-28 cs.CV 新提交 83%

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

迈向视觉语言导航的双脑最小充分表示

Yihao Wu, Chenyi Xu, Liqi Yan, Chenhuan Cai, Geyong Min, Bin Lin, Fangli Guan, Jianhui Zhang, Pan Li

机构 * Hangzhou Dianzi University(杭州电子科技大学) University of Zurich(苏黎世大学) University of Exeter(埃克塞特大学)

专题命中 具身与机器人 :world model(abstract);world-model(abstract);world model(abstract);world-model(abstract)

AI总结 针对视觉语言导航中存在的问题,提出BrainNav框架,包含逻辑锚定模型等三个组件,通过将语义意图与空间感知对齐,提升智能体在复杂任务中的表现,在实验中相比之前最优方法有改进,为鲁棒的视觉语言导航提供有效基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12090 2026-05-13 cs.RO cs.CL cs.CV 83%

World Action Models: The Next Frontier in Embodied AI

世界动作模型:具身AI的下一个前沿

Siyin Wang, Junhao Shi, Zhaoyang Fu, Xinzhe He, Feihong Liu, Chenchen Yang, Yikang Zhou, Zhaoye Fei, Jingjing Gong, Jinlan Fu, Mike Zheng Shou, Xuanjing Huang, Xipeng Qiu, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) National University of Singapore(新加坡国立大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文探讨了具身AI中的世界动作模型(WAMs),通过整合环境动态预测模型,统一预测状态建模与动作生成,提出结构化分类体系并分析数据生态,为该领域提供系统性综述。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26255 2026-03-17 cs.AI cs.CV cs.LG cs.RO 83%

ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning

ExoPredicator:为机器人规划学习动态世界的抽象模型

Yichao Liang, Dat Nguyen, Cambridge Yang, Tianyang Li, Joshua B. Tenenbaum, Carl Edward Rasmussen, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出ExoPredicator框架,通过联合学习符号状态表示和因果过程,解决长周期具身规划中的外源性过程建模问题,实现对复杂任务的泛化能力。

Comments ICLR 2026. The last two authors contributed equally in co-advising

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10675 2026-01-07 cs.RO cs.AI cs.CV cs.LG 83%

Evaluating Gemini Robotics Policies in a Veo World Simulator

在Veo世界模拟器中评估Gemini机器人策略

Gemini Robotics Team, Krzysztof Choromanski, Coline Devin, Yilun Du, Debidatta Dwibedi, Ruiqi Gao, Abhishek Jindal, Thomas Kipf, Sean Kirmani, Isabel Leal, Fangchen Liu, Anirudha Majumdar, Andrew Marmon, Carolina Parada, Yulia Rubanova, Dhruv Shah, Vikas Sindhwani, Jie Tan, Fei Xia, Ted Xiao, Sherry Yang, Wenhao Yu, Allan Zhou

机构 * Gemini Robotics Team(Gemini机器人团队) Google DeepMind(谷歌DeepMind)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本研究提出基于Veo视频模型的生成式评估系统,用于在Veo世界模拟器中全面评估Gemini机器人策略的性能,包括名义性能、分布外泛化及安全约束测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14198 2025-06-18 cs.RO cs.CV cs.LG 83%

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

Jeremy A. Collins, Loránd Cheng, Kunal Aneja, Albert Wilcox, Benjamin Joffe, Animesh Garg

机构 * Georgia Tech(佐治亚理工学院) Georgia Tech Research Institute(佐治亚理工研究 institute)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03227 2022-07-14 cs.RO cs.AI cs.CL cs.CV cs.LG 83%

CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

Oier Mees, Lukas Hermann, Erick Rosete-Beas, Wolfram Burgard

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Accepted for publication at IEEE Robotics and Automation Letters (RAL). Code, models and dataset available at http://calvin.cs.uni-freiburg.de

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25880 2026-06-25 cs.CV 新提交 83%

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning

USS: 面向具身视觉跟踪的统一空间-语义提示与潜在动力学学习

Yuchen Xie, Xinyu Zhou, Kuangji Zuo, Yanshuo Lu, Fengrui Huang, Boyu Ma, Jianfei Yang

机构 * Nanyang Technological University(南洋理工大学)

专题命中 具身与机器人 :latent dynamics(title,abstract);world model(abstract);world model(abstract);分类 cs.CV

AI总结 提出统一空间-语义提示范式替代纯文本目标指示,设计端到端框架USS支持多种提示类型,通过潜在世界模型提升时序鲁棒性,在真实机器人实验中优于纯文本方法。

详情

展开后加载摘要…

URL PDF HTML 收藏