Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
鲁棒的音频视觉目标说话人提取与情感感知多注册融合
Zhan Jin, Bang Zeng, Peijun Yang, Jiarong Du, Wei Ju, Yao Tian, Juan Liu, Ming Li
机构
*
School of Computer Science, Wuhan University, Wuhan, China(1 武汉大学计算机学院,武汉,中国)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(2 香港中文大学(深圳)人工智能学院,中国)
;
School of Artificial Intelligence, Wuhan University, Wuhan, China(3 武汉大学人工智能学院,武汉,中国)
;
School of Cyber Science and Engineering, Wuhan University, Wuhan, China(4 武汉大学网络科学与工程学院,武汉,中国)
;
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China(5 香港中文大学(深圳)数字创新研究中心,中国)
;
AI Center, OPPO, Beijing, China(6 OPPO人工智能中心,北京,中国)
Yunsheng Wang, Yuntao Shou, Yilong Tan, Wei Ai, Tao Meng, Keqin Li
机构
*
College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南林业科技大学)
;
Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)
机构
*
Google Mountain View CA USA(谷歌山景城)
;
University of Michigan Ann Arbor MI USA(密歇根大学安 Arbor分校)
;
Columbia University New York NY USA(哥伦比亚大学)
;
Google San Francisco CA USA(谷歌旧金山)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
高效音频-视觉语音分离与离散唇语语义及多尺度全局-局部注意力
Kai Li, Kejun Gao, Xiaolin Hu
机构
*
Department of Computer Science and Technology, Institute for AI, BNRist, Tsinghua University, Beijing(计算机科学与技术系,人工智能研究院,BNRist,清华大学,北京)
;
IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing(IDG/麦克戈文脑研究 institute,清华大学,北京)
;
Chinese Institute for Brain Research (CIBR), Beijing(中国脑科学研究院(CIBR),北京)