Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
基于多模态大基础模型的歌唱音色受欢迎程度评估
机构 * Zhejiang University(浙江大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Hong Kong University of Science and Technology(香港科学与技术大学) ; University of California, Berkeley(加州大学伯克利分校) ; University of Manchester(曼彻斯特大学) ; Innovation Center of Yangtze River Delta, Zhejiang University(长江三角洲创新中心,浙江大学)
AI总结 本文提出基于多模态大基础模型的歌唱音色受欢迎程度评估方法,通过引入Sing-MD数据集、VocalVerse架构和H-TPR基准,实现无参考、多维度的歌唱评估。
Comments Accepted to ACMMM 2025 oral
Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (ACMMM 2025), Pages 12227-12236