DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
DepthCache: 一种基于深度的无训练视觉标记融合方法,用于视觉-语言-动作模型推理
机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) ; School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)
AI总结 DepthCache通过利用深度作为结构先验,实现无训练的视觉标记融合,提升视觉-语言-动作模型推理速度,同时保持较高的任务成功率。
Comments 8 pages, 6 figures