USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots
USIM 和 U0:面向通用水下机器人的视觉-语言-动作数据集与模型
机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) ; Baidu Inc.(百度公司) ; The School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract,abstract_cn);分类 cs.RO
AI总结 提出一个统一框架,通过数据合成管道构建模拟数据集 USIM 和视觉-语言-动作模型 U0,实现水下机器人从避障导航到三维移动操作的多任务执行,在离线动作预测误差和在线成功率上达到最优。
Comments Project Page: https://vincentgu2000.github.io/u0project/