Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
Cheers:解耦补丁细节与语义表示以实现统一的多模态理解和生成
机构 * Tsinghua University(清华大学) ; Xi’an Jiaotong University(西安交通大学) ; University of Chinese Academy of Sciences(中国科学院大学)
AI总结 Cheers通过解耦补丁细节与语义表示,实现统一的多模态理解和生成,提升图像生成的保真度,并在多个基准测试中表现优异,同时实现4倍的token压缩效率。
Comments 17 pages, 5 figures