Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
多模态流:嵌入空间中语言与视觉的统一流建模
机构 * Huazhong University of Science and Technology(华中科技大学) ; Beijing Jiaotong University(北京交通大学) ; Horizon Robotics(地平线机器人)
AI总结 多模态流提出一种完全连续的生成模型,通过共享分块因果流主干统一语言和视觉,在多个规模上持续预训练,以较少数据达到竞争性能,确立了新的连续多模态范式。
Comments 18 pages, 5 figures, 10 tables. Code and model: https://github.com/hustvl/Multimodal-Flow