From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
从推理到像素:统一多模态模型中对齐差距的基准测试
机构 * University of California San Diego(加利福尼亚大学圣迭戈分校) ; University of Southern California(南加利福尼亚大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Carnegie Mellon University(卡内基梅隆大学)
AI总结 本文通过UReason基准测试,探讨统一多模态模型中模态对齐问题,发现去上下文生成在图像生成任务中表现更优,揭示了文本推理与生成图像之间存在对齐差距。
Comments Project page: https://ureason.github.io