VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence
VLT:用于工业智能的视觉-语言-时间序列多模态基础模型
机构 * School of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) ; Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) ; State Key Laboratory of Intelligent Manufacturing System Technology(智能制造系统技术国家重点实验室) ; School of Computer Science and Artificial Intelligence, Zhengzhou University(郑州大学计算机科学与人工智能学院)
专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 针对工业时间序列单模态建模局限及连接时间序列与文本语义的挑战,提出VLT多模态基础模型,通过设计Time-MoE等机制联合建模多种模态,经实验验证其在多种复杂设置下优于现有方法,提升了鲁棒性和泛化能力。
Comments 18 pages, 13 figures, and 13 tables, including supplementary material. Haiteng Wang and Jingheng Yan contributed equally to this work