arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大厂专区

2026-03-12 至 2026-03-12 共收录 1
2509.12583 2026-03-12 eess.AS cs.SD

Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

鲁棒的音频视觉目标说话人提取与情感感知多注册融合

Zhan Jin, Bang Zeng, Peijun Yang, Jiarong Du, Wei Ju, Yao Tian, Juan Liu, Ming Li

机构 * School of Computer Science, Wuhan University, Wuhan, China(1 武汉大学计算机学院,武汉,中国) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(2 香港中文大学(深圳)人工智能学院,中国) School of Artificial Intelligence, Wuhan University, Wuhan, China(3 武汉大学人工智能学院,武汉,中国) School of Cyber Science and Engineering, Wuhan University, Wuhan, China(4 武汉大学网络科学与工程学院,武汉,中国) Digital Innovation Research Center, Duke Kunshan University, Kunshan, China(5 香港中文大学(深圳)数字创新研究中心,中国) AI Center, OPPO, Beijing, China(6 OPPO人工智能中心,北京,中国)

AI总结 本文提出一种鲁棒的音频视觉目标说话人提取方法,通过情感感知的多注册融合技术,在模态缺失情况下提升提取性能和鲁棒性。

Comments submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏