Learning to Expand Images for Efficient Visual Autoregressive Modeling
学习扩展图像以实现高效的视觉自回归建模
机构 * University of Electronic Science and Technology of China(电子科技大学) ; School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) ; Xidian University(西安电子科技大学) ; SenseTime Research(商汤科技研究院) ; Shanghai Jiao Tong University(上海交通大学)
AI总结 本文提出EAR模型,通过中心向外的生成方式提升视觉自回归模型的效率与生成质量,实现保真度与效率的最佳平衡。
Comments 16 pages, 18 figures, includes appendix with additional visualizations, submitted as arXiv preprint