DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
DepthCache: 一种基于深度的无训练视觉标记融合方法,用于视觉-语言-动作模型推理
Yuquan Li, Lianjie Ma, Han Ding, Lijun Zhu
机构
*
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)
机构
*
NVIDIA
;
The Chinese University of Hong Kong(香港中文大学)
;
Sung Kyun Kwan University(全州大学)
;
Wenzhou Medical University(温州医学院)
;
National University of Singapore(新加坡国立大学)
;
Ruijin Hospital(瑞金医院)
专题命中
VLA模型
:vision language action(abstract);VLA(abstract);分类 cs.RO、cs.CV