Towards Scalable Pre-training of Visual Tokenizers for Generation
面向生成任务的视觉分词器可扩展预训练
机构 * Huazhong University of Science and Technology(华中科技大学) ; MiniMax
AI总结 VTP通过联合优化图像-文本对比、自监督和重建损失,提升视觉分词器的生成性能和扩展性。
Comments Our pre-trained models are available at https://github.com/MiniMax-AI/VTP