vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
vla.cpp:视觉-语言-动作模型的统一推理运行时
机构 * VinRobotics ; Center for AI Research, VinUniversity(VinUniversity 人工智能研究中心) ; Intelligent Autonomous Systems, TU Darmstadt(达姆施塔特工业大学智能自主系统) ; Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究学院) ; University of Stuttgart(斯图加特大学) ; German Research Center for Artificial Intelligence(德国人工智能研究中心)
专题命中 VLA模型 :VLA(title,title_cn);vision-language-action(title,abstract);action model(title);分类 cs.RO、cs.AI、cs.LG
AI总结 提出vla.cpp,基于llama.cpp的便携C++推理运行时,支持多种VLA架构,在LIBERO-Object上接近SOTA性能,内存仅1.3 GiB,并实现跨硬件部署。
Comments 17 pages, 3 figures, 12 tables