Le MuMo JEPA: Multi-Modal Self-Supervised Representation Learning with Learnable Fusion Tokens
Le MuMo JEPA:多模态自监督表示学习中的可学习融合标记
机构 * IDLab, Department of Information Technology, Ghent University - imec(根特大学-imec信息技术系IDLab)
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract,comments);multimodal(abstract);分类 cs.CV
AI总结 本文提出Le MuMo JEPA框架,通过学习融合标记实现多模态统一表示学习,实验显示其在性能与效率之间取得最佳平衡,尤其在CenterNet检测和密集深度估计中表现优异。
Comments 14 pages, 4 figures, supplementary material. Accepted at the CVPR 2026 Workshop on Unified Robotic Vision with Cross-Modal Sensing and Alignment (URVIS)