Vehicle-centric Perception via Multimodal Structured Pre-training
基于多模态结构预训练的车辆感知
机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province(安徽省信息材料与智能感知实验室) ; Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室) ; the School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) ; School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 本文提出VehicleMAE-V2,通过多模态结构先验知识提升车辆感知的预训练能力,采用SMM、CRM和SRM模块增强模型对车辆结构和语义的理解。
Comments Journal extension of VehicleMAE (AAAI 2024)