LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
LVDrive: 基于潜在视觉表征的视觉-语言-动作自动驾驶模型
机构 * The Hong Kong University of Science and Technology(香港科技大学) ; Xiaomi EV(小米电动车)
专题命中 端到端驾驶 :autonomous driving(title,abstract);分类 cs.CV、cs.AI
AI总结 本文提出LVDrive,一种增强视觉-语言-动作能力的自动驾驶模型,通过引入未来场景预测任务,在高维潜在空间中学习语义丰富的场景表示,从而提升闭环驾驶性能。