D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
机构 * Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系) ; Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所) ; Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中 其他安全 :safety(abstract);分类 cs.AI
AI总结 本文针对自动驾驶中多模态大语言模型,提出D3VL框架,整合2D和3D时间序列数据,回答交通场景相关问题,在KITTI问答数据集上性能提升11%,并引入Waymo QA数据集扩展以评估模型在多样驾驶条件下处理3D和时间序列数据的能力。
Comments Accepted to IEEE IV 2026