S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving
S平方-VLA:自动驾驶视觉-语言-动作模型中语义与空间流的解耦
机构 * School of Mechanical and Electronic Engineering, Wuhan University of Technology(武汉理工大学机电工程学院) ; Intelligent Transportation Systems Research Center, Wuhan University of Technology(武汉理工大学智能交通系统研究中心) ; School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
专题命中 规划控制 :autonomous driving(title,abstract);trajectory planning(abstract);分类 cs.RO
AI总结 研究针对自动驾驶中视觉语言模型生成低级控制动作的局限,提出S平方-VLA解耦语义和空间流,语义流用于意图推理,空间流保留空间特征并赋予先验,双流规划适配器融合二者,在基准测试中取得新的最先进水平,优于基线。