Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection
弦外之音:ASR嵌入与LLM增强语言学联合学习用于痴呆检测
机构 * Division of Communication and Media, Ewha Womans University, South Korea(韩国梨花女子大学传播与媒体学院) ; NAVER Cloud, South Korea(韩国NAVER Cloud)
专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)
AI总结 提出多模态框架,结合Whisper的声学表示与LLM提取的语言特征,通过门控融合网络实现痴呆检测,在ADReSS和ADReSSo上F1分别达89.47%和90.14%。
Comments Accepted at INTERSPEECH 2026