TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
TaigiSpeech: 一个低资源真实世界语音意图数据集及基于可扩展野外数据挖掘的初步结果
机构 * Massachusetts Institute of Technology, USA(麻省理工学院) ; National Taiwan University, Taipei, Taiwan(国立台湾大学) ; National Taiwan University Artificial Intelligence Center of Research Excellence, Taipei, Taiwan(国立台湾大学人工智能研究中心) ; Academia Sinica, Taiwan(台湾“中央”研究院) ; National Yang Ming Chiao Tung University, Taiwan(阳明交通大学) ; Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California, USA(信号分析与解释实验室(SAIL),南加州大学)
专题命中 评测与基准 :LLM(summary_cn,abstract);分类 cs.CL、cs.LG
AI总结 针对低资源台语,构建包含21位老年人3000条话语的语音意图数据集,并探索关键词匹配与LLM伪标注、音视频框架两种数据挖掘策略,以解决标注数据稀缺问题。
Comments Interspeech 2026 long paper