Teaching Vision-Language-Action Models What to See and Where to Look
教视觉-语言-动作模型看什么和看哪里
机构 * School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院) ; National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院) ; Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) ; DiDi(滴滴出行) ; State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室) ; School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) ; School of Cyber Science and Technology, Beihang University(北京航空航天大学网络安全科学与技术学院) ; School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
专题命中 VLA模型 :VLA(summary_cn,abstract);vision-language-action(title,abstract);action model(title);分类 cs.CV
AI总结 提出DriveTeach-VLA框架,通过驾驶感知视觉蒸馏和2D轨迹引导提示,教VLA模型关注驾驶相关区域,实现端到端自动驾驶轨迹预测,在NAVSIM和nuScenes上达到最优性能。
Comments The paper has been accepted by ECCV 2026