S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
S-VAM:通过自蒸馏几何和语义前瞻性构建快捷视频动作模型
机构 * The Hong Kong University of Science ; Technology (Guangzhou) Huawei Foundation Model Department
专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.CV
AI总结 S-VAM通过自蒸馏策略实现单次前向传递生成几何和语义表示,提升动作预测效率,实验证明在复杂环境中优于现有方法。