How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
如何合成高质量的预训练数据?提示设计、生成模型和源数据的系统研究
机构 * Hugging Face ; University of Sheffield(谢菲尔德大学) ; National Institute of Informatics(日本信息处理研究所)
AI总结 本文通过大规模实验探讨提示设计、生成模型和源数据对合成预训练数据质量的影响,发现结构化输出格式优于现有方法,且生成模型参数超过1B无额外收益,提出开源数据集FinePhrase。
Comments Accepted to COLM 2026