Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
面向表达性与忠实性的音频到图像生成:一个统一的多模态数据集与合成框架
机构 * University of Science and Technology of China(中国科学技术大学) ; China Telecom (TeleAI)(中国电信(电信人工智能研究院)) ; Northwest Polytechnical University(西北工业大学)
专题命中 多模态生成 :multimodal(title);cross-modal(abstract);audio-visual(abstract);分类 cs.CV
AI总结 针对音频到图像生成受限于传统数据集的问题,提出A2I-Set数据集与AudioCanvas模型,实现了更优的跨模态对齐与视觉表达性。
Comments 23 pages, 16 figures