Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings
面向开放集说话人属性预测的关键词附加大语言模型嵌入
机构 * Department of Intelligence and Information, Seoul National University(首尔大学情报信息学系) ; Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔大学人工智能跨学科项目) ; Artificial Intelligence Institute, Seoul National University(首尔大学人工智能研究所)
专题命中 音频语音多模态 :cross-modal(abstract)
AI总结 提出利用大语言模型嵌入进行开放集说话人属性预测,通过关键词附加策略和top-k负损失,在LibriTTS-P上超越闭集基准并泛化到未见同义词。
Comments This paper has been accepted to Interspeech 2026