VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
VLA-4D: 将4D意识嵌入视觉-语言-动作模型中以实现时空一致的机器人操作
机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) ; School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
AI总结 VLA-4D通过4D意识增强视觉-语言-动作模型,实现时空一致的机器人操作,提升动作执行的空间平滑性和时间一致性。