arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

共收录 9088 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 9088 篇

2601.21712 2026-02-04 cs.RO 88%

CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation

CoFreeVLA:通过视觉-语言-动作模型和风险估计实现碰撞-free双臂操作

Xuanran Zhai, Binkai Ou, Qiaojun Yu, Ce Hao, Yaohua Liu

专题命中 VLA模型 :vision-language-action(title);action model(title);vision language action(abstract);VLA(abstract)

AI总结 CoFreeVLA通过视觉-语言-动作模型和风险估计实现双臂操作的安全性提升,有效减少自我碰撞并提高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21602 2026-02-04 cs.RO 88%

AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation

AIR-VLA:面向空中操作的视觉-语言-动作系统

Jianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian

机构 * The Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) School of Information and Intelligent Engineering, University of Sanya(三亚大学信息与智能工程学院) School of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) State Key Lab of Intelligent Transportation Systems, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);分类 cs.RO

AI总结 AIR-VLA提出首个针对空中操作的视觉-语言-动作系统,通过构建仿真环境和多模态数据集,评估主流模型并揭示其在无人机移动、机械臂控制和高层规划中的能力和限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00780 2026-02-03 cs.AI 88%

Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

环境感知自适应剪枝与交错推断调度用于视觉-语言-动作模型

Yuting Huang, Leilei Ding, Zhipeng Tang, Zenghuan Zhu, Jiajun Deng, Xinrui Lin, Shuo Liu, Haojie Ren, Jianmin Ji, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.AI

AI总结 EcoVLA通过环境感知自适应剪枝和交错推断调度,实现视觉-语言-动作模型的高效参数稀疏化,提升推理速度并保持高成功率。

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00686 2026-02-03 cs.RO 88%

Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching

通过自适应视觉令牌缓存学习加速视觉-语言-动作模型

Yujie Wei, Jiahan Fan, Jiyu Guo, Ruichen Zhen, Rui Shao, Xiu Su, Zeke Xie, Shuo Yang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Meituan Academy of Robotics Shenzhen, Meituan(美团机器人深圳研究院) Central South University(中南大学) HKUST(GZ)(香港科技大学(广州))

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 本文提出通过自适应视觉令牌缓存学习加速VLA模型,提升推理效率并提高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00500 2026-02-03 cs.RO 88%

Inject Once Survive Later: Backdooring Vision-Language-Action Models to Persist Through Downstream Fine-tuning

一次注入,后期存活:向视觉-语言-动作模型注入后门以在下游微调中持续存在

Jianyi Zhou, Yujie Wei, Ruichen Zhen, Bo Zhao, Xiaobo Xia, Rui Shao, Xiu Su, Shuo Yang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Harbin Institute of Technology(哈尔滨工业大学) Meituan Academy of Robotics Shenzhen, Meituan(美团机器人深圳研究院) Shanghai Jiaotong University(上海交通大学) National University of Singapore(新加坡国立大学) Central South University(中南大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 本文提出INFUSE框架,首次实现对VLA模型的后门攻击,使其在用户微调后仍能保持有效性,显著提升攻击成功率并保持任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25889 2026-01-30 cs.LG 88%

$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

$π_\ exttt{RL}$: 流基于视觉-语言-动作模型的在线强化学习微调

Kang Chen, Zhihao Liu, Tonghe Zhang, Zhen Guo, Si Xu, Hao Lin, Hongzhi Zang, Xiang Li, Quanlu Zhang, Zhaofei Yu, Guoliang Fan, Tiejun Huang, Yu Wang, Chao Yu

机构 * Tsinghua University(清华大学) Peking University(北京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Carnegie Mellon University(卡内基梅隆大学) Infinigence AI Zhongguancun Academy(中关村学院)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.LG

AI总结 本文提出$π_\ exttt{RL}$方法,通过流噪声和流SDE技术,解决大规模流基于VLA模型中强化学习微调的挑战,提升模型在分布内和分布外任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08325 2026-01-14 cs.RO 88%

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

ActiveVLA: 向视觉-语言-动作模型注入主动感知以实现精确的3D机器人操控

Zhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Nanyang Technological University(南洋理工大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 ActiveVLA通过引入主动感知能力,提升机器人在复杂环境中的高精度3D操控性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03044 2026-01-07 cs.RO 88%

SOP: A Scalable Online Post-Training System for Vision-Language-Action Models

SOP:一种可扩展的在线后训练系统用于视觉-语言-动作模型

Mingjie Pan, Siyuan Feng, Qinglin Zhang, Xinchen Li, Jianheng Song, Chendi Qu, Yi Wang, Chuankang Li, Ziyu Xiong, Zhi Chen, Yi Liu, Jianlan Luo

机构 * Agibot Research(Agibot研究机构) Shanghai Innovation Institute(上海创新研究院)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 SOP系统通过在线分布式后训练提升视觉-语言-动作模型在现实世界中的性能,实现高效可靠的多任务适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11315 2025-12-15 cs.LG 88%

Benchmarking the Generality of Vision-Language-Action Models

对视觉-语言-动作模型通用性的基准测试

Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka, Mridul Sharma, Helen Lu, Sean Rivera, Aryan Khurana, Hangliang Ren, Yangyue Wang

机构 * Manifold Research Metarch AI Georgia Tech(佐治亚理工学院) Tufts University(塔夫茨大学) Northeastern University(东北大学) Birla Institute of Technology and Science, Pilani(比拉理工学院,帕利尼) Institute for Research and Innovation in Intelligent Systems (IRIIS)(智能系统研究与创新研究所)

专题命中 VLA模型 :action model(title,abstract);vision-language-action(title);vision language action(abstract);分类 cs.LG

AI总结 本文提出MultiNet v1.0基准,评估视觉-语言-动作模型在六个基础能力领域的跨领域泛化能力,发现现有模型在未见领域和模态转移时表现显著退化。

Comments 23 pages, 7 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09927 2025-12-11 cs.RO 88%

Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models

令牌扩展-合并:面向视觉-语言-动作模型的免训练 令牌压缩

Yifan Ye, Jiaqi Ma, Jun Cen, Zhihe Lu

机构 * College of Science and Engineering, Hamad Bin Khalifa University(1 科学与工程学院,哈马德·本·卡西姆大学) Mohamed bin Zayed University of Artificial Intelligence(2 摩萨·本·扎耶德人工智能大学) College of Computer Science and Technology, Zhejiang University(3 计算机科学与技术学院,浙江大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 TEAM-VLA通过动态令牌扩展与合并机制,实现无需训练的视觉-语言-动作模型高效推理,提升速度并保持任务性能。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09619 2025-12-11 cs.RO 88%

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

GLaD:面向视觉-语言-动作模型的几何潜在蒸馏

Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu, Yan Yan, Xiaodan Liang, Ivan Laptev, Xiaojun Chang

机构 * MBZUAI University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 GLaD通过引入几何意识的预训练机制,提升了视觉-语言-动作模型的空间推理和策略泛化能力,无需依赖深度传感器或3D标注。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07582 2025-12-09 cs.RO 88%

See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations

一次观看,随后行动:基于单次视频示范的任务学习视觉-语言-行动模型

Guangyan Chen, Meiling Wang, Qi Shao, Zichen Zhou, Weixin Mao, Te Cui, Minzhao Zhu, Yinan Deng, Luojie Yang, Zhanqi Zhang, Yi Yang, Hua Chen, Yufeng Yue

机构 * Beijing Institute of Technology(北京理工大学) LimX Dynamics

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 ViVLA通过单次视频示范高效学习机器人操控任务,实现跨任务和跨身体的显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04446 2025-12-05 cs.RO cs.SY eess.SY 88%

Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops

面向选择性机器人拆解的视觉-语言-动作模型:以从台式机中提取关键部件的案例研究

Chang Liu, Sibo Tian, Sara Behdad, Xiao Liang, Minghui Zheng

机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University(J. Mike Walker ’66 机械工程系,德克萨斯A&M大学) Engineering School of Sustainable Infrastructure & Environment, University of Florida(可持续基础设施与环境工程学院,佛罗里达大学) Zachry Department of Civil and Environmental Engineering, Texas A&M University(土木与环境工程系,德克萨斯A&M大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 本文研究了视觉-语言-动作模型在机器人选择性拆解中的应用,通过定制数据集微调两种VLA方法,发现其在复杂拆解任务中存在局限,但结合规则控制器的混合策略能有效完成拆解任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04018 2025-12-04 cs.RO 88%

FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction

FPC-VLA:一个带有监督器的视觉-语言-动作框架,用于故障预测和纠正

Yifan Yang, Zhixiang Duan, Tianshi Xie, Fuyu Cao, Pinxi Shen, Peili Song, Piaopiao Jin, Guokang Sun, Shaoqing Xu, Yangwei You, Jingtai Liu

机构 * The Institute of Robotics and Automatic Information System(机器人与自动信息系统研究所) Tianjin Key Laboratory of Intelligent Robotics(智能机器人天津重点实验室) TBI Center, Nankai University, Tianjin 300350, China(南开大学天津中心) Faculty of Robot Science and Engineering, Northeastern University, Shenyang 110819, China(机器人科学与工程学院) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence(人工智能中心) Department of Electromechanical Engineering, University of Macau, Macau SAR, China(机电工程系,澳门大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);分类 cs.RO

AI总结 FPC-VLA 提出了一种双模型框架,结合视觉-语言-动作模块与监督器,用于预测和纠正机器人操作中的故障,提升了自主系统的可靠性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23463 2025-11-24 cs.CV 88%

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

OpenDriveVLA: 向端到端自动驾驶迈进的大型视觉语言动作模型

Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, Volker Tresp, Alois Knoll

机构 * Technical University of Munich(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 VLA模型 :vision language action(title,abstract);action model(title,abstract);分类 cs.CV

AI总结 OpenDriveVLA基于开源大语言模型,通过多模态输入和分层视觉语言对齐,实现端到端自动驾驶中的高精度轨迹规划和驾驶任务回答。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16166 2025-11-21 cs.CV 88%

EvoVLA: Self-Evolving Vision-Language-Action Model

EvoVLA:自进化视觉-语言-动作模型

Zeting Liu, Zida Yang, Zeyu Zhang, Hao Tang

机构 * Peking University(北京大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.CV

AI总结 EvoVLA通过自监督框架解决VLA模型的阶段幻觉问题,提升任务成功率和样本效率,实现有效的仿真到现实转移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13757 2025-11-07 cs.CV 88%

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, Jiaqi Ma

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.CV

Comments NeurIPS 2025; Website link:https://autovla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13174 2025-11-06 cs.CV 88%

Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models

Hao Cheng, Erjia Xiao, Yichi Wang, Chengyuan Yu, Mengshu Sun, Qiang Zhang, Jiahang Cao, Yijie Guo, Ning Liu, Kaidi Xu, Jize Zhang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) City University of Hong Kong(香港城市大学) Beijing University of Technology(北京理工大学) Duke University(杜克大学) X-Humanoid Project(X-Humanoid 项目)

专题命中 VLA模型 :vision language action(title,abstract);action model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01224 2025-11-04 cs.RO 88%

Embodiment Transfer Learning for Vision-Language-Action Models

Chengmeng Li, Yaxin Peng

机构 * Shanghai University(上海大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25713 2025-10-30 cs.RO 88%

Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models

Boshi An, Chenyu Yang, Robert Katzschmann

机构 * Soft Robotics Lab, ETHz, Switzerland(苏黎世联邦理工学院软机器人实验室)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18337 2025-10-24 cs.RO 88%

MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning

Wenhui Huang, Changhe Chen, Han Qi, Chen Lv, Yilun Du, Heng Yang

机构 * Harvard University(哈佛大学) University of Michigan(密歇根大学) Nanyang Technological University(南洋理工大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12276 2025-10-20 cs.RO 88%

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tsinghua University(清华大学) Westlake University(西湖大学) Zhejiang University(浙江大学) South China University of Technology(华南理工大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13375 2025-10-16 cs.CV 88%

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Zhuoguang Chen, Tao Jiang, Hang Zhao

机构 * IIIS, Tsinghua University(清华大学智能技术学部)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10975 2025-10-15 cs.RO 88%

RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model

Mingtong Dai, Lingbo Liu, Yongjie Bai, Yang Liu, Zhouxia Wang, Rui SU, Chunjie Chen, Liang Lin, Xinyu Wu

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Peng Cheng Laboratory(鹏城实验室) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Shanghai AI Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) X-Era AI Lab(X-Era人工智能实验室)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01424 2025-10-14 cs.RO 88%

TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control

Zhenyang Liu, Yongchong Gu, Sixiao Zheng, Yanwei Fu, Xiangyang Xue, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18269 2025-10-08 cs.RO 88%

FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models

Zhide Zhong, Haodong Yan, Junfeng Li, Xiangchen Liu, Xin Gong, Tianran Zhang, Wenxuan Song, Jiayi Chen, Xinhu Zheng, Hesheng Wang, Haoang Li

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19870 2025-09-25 cs.CV 88%

FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models

Xin Wang, Jie Li, Zejia Weng, Yixu Wang, Yifeng Gao, Tianyu Pang, Chao Du, Yan Teng, Yingchun Wang, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室) Sea AI Lab(Sea人工智能实验室)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14138 2025-09-18 cs.RO 88%

SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model

Ran Yang, Zijian An, Lifeng ZHou, Yiming Feng

机构 * Virginia Seafood Agricultural Research and Extension Center, and Department of Biological Systems Engineering, Virginia Tech(弗吉尼亚海产品农业研究中心及生物系统工程系,弗吉尼亚理工学院) Department of Electrical and Computer Engineering, Drexel University(电气与计算机工程系,德雷塞尔大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

Comments 8 pages, 9 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09071 2025-08-14 cs.RO 88%

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Lin Sun, Bin Xie, Yingfei Liu, Hao Shi, Tiancai Wang, Jiale Cao

机构 * Tianjin University(天津大学) Dexmal Tsinghua University(清华大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

Comments The project is visible at https://linsun449.github.io/GeoVLA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06547 2025-08-12 cs.RO 88%

A tutorial note on collecting simulated data for vision-language-action models

Heran Wu, Zirun Zhou, Jingfeng Zhang

机构 * School of Computer Science, The University of Auckland(计算机科学学院,奥克兰大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

Comments This is a tutorial note for educational purposes

详情

展开后加载摘要…

URL PDF HTML 收藏