Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Crab$^{+}$: 一种可扩展且统一的音频视觉场景理解模型,具有显式合作
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院) ; Institute of Artificial Intelligence of China Telecom (TeleAI)(中国电信人工智能研究院) ; AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频业务单元AI技术中心) ; Hefei University of Technology(合肥工业大学) ; The University of Hong Kong(香港大学)
AI总结 Crab$^{+}$通过显式合作解决音频视觉任务异质性问题,实现更广泛的任务覆盖和优于单任务模型的性能表现。