FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
FlowInOne:将多模态生成统一为图像输入、图像输出的流匹配
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Central South University(中南大学) ; National University of Singapore(新加坡国立大学) ; University of Science and Technology of China(中国科学技术大学) ; Microsoft(微软)
专题命中 扩散模型 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 本文提出FlowInOne框架,将多模态生成统一为纯视觉流匹配,通过视觉提示消除跨模态对齐瓶颈,实现文本到图像生成、布局引导编辑和视觉指令遵循的统一。
Comments Accepted by ECCV 2026. 38 Pages, 21 Figures, 12 Tables