VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
VFEM: 视觉特征赋能的多变量时间序列预测与跨模态融合
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) ; Pengcheng Laboratory(鹏城实验室) ; Ant Group(蚂蚁集团) ; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) ; University of Pennsylvania(宾夕法尼亚大学)
专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI
AI总结 提出VFEM模型,利用预训练大视觉模型通过跨模态注意力融合视觉与时间特征,仅训练7.45%参数即可捕捉跨变量依赖,提升多变量时间序列预测性能。