COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
COMET:音频-文本多模态对比嵌入中模态间隙的概念空间剖析
机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) ; Centre for Vision, Speech, and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS
AI总结 提出COMET框架,通过PLS-SVD分解揭示CLAP模型中模态间隙主要由少数共享概念轴贡献,并基于谱截断方法无训练地缓解间隙,实现零样本音频字幕接近全监督性能。