RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
RA-SSU:迈向细粒度音频视觉学习的区域感知声音源理解
Muyi Sun, Yixuan Wang, Hong Wang, Chen Su, Man Zhang, Xingqun Qi, Qi Li, Zhenan Sun
机构
*
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(香港科技大学跨学科研究院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)