When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
当微调失效且泛化时:数据多样性与混合训练在基于大语言模型的文本到语音系统中的作用
机构 * Sprinklr AI
专题命中 指令微调 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 本文研究了LoRA微调在语音克隆任务中的有效性,发现其在语音质量和信噪比提升方面优于基模型,尤其在数据多样性支持下表现更佳。
Comments We finetune the Qwen 0.5B backbone in an LLM TTS with LoRA to raise MOS speaker similarity and SNR. It works best with diverse training audio with uniform data it can amplify noise so tune decoding and use GGUF quantization for low latency stable quality