M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition
M2S-AVSR:面向鲁棒视听语音识别的模态感知多视角自监督表示
机构 * School of Artificial Intelligence and the School of Computer Science, Wuhan University, China(人工智能学院和计算机科学学院,武汉大学,中国) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(人工智能学院,香港中文大学(深圳),中国) ; School of Artificial Intelligence, Wuhan University, China(人工智能学院,武汉大学,中国)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 eess.AS
AI总结 提出一种模态感知多视角自监督表示框架(M2S-AVSR),通过多视角编码学习视角不变视觉语音表示,并利用模态感知模块进行细粒度融合,以应对视角变化、音频失真和视觉遮挡等挑战,在多个基准上取得最优性能。
Comments submitted to IEEE Transactions on Audio, Speech, and Language Processing