VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
VLADriver-RAG:基于检索增强的视觉-语言-动作模型用于自动驾驶
机构 * College of Automotive Engineering(汽车工程学院) ; The National Key Laboratory of Automotive Chassis Integration and Bionics(汽车底盘集成与生物力学国家级重点实验室) ; ReeFocus AI Technology(ReeFocus人工智能技术)
专题命中 多模态RAG :RAG(title,title_cn);retrieval-augmented generation(abstract);分类 cs.AI
AI总结 本文提出VLADriver-RAG,通过构建结构化的历史知识,提升自动驾驶中长尾场景的泛化能力,采用视觉到场景机制和场景对齐嵌入模型,结合查询驱动的VLA骨干网络,实现高精度轨迹合成。