How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
光学流和文本提示如何协作以辅助音频视觉语义分割?
机构 * Hong Kong Baptist University(香港 Baptist 大学) ; Guangdong Provincial/Zhuhai Key Laboratory IRADS(广东省级/珠海关键实验室 IRADS) ; Department of Computer Science, Beijing Normal-Hong Kong Baptist University, Zhuhai, China(计算机科学系,北京师范大学-香港 Baptist 大学,珠海,中国) ; Peking University, Shenzhen Graduate School, Shenzhen, China(北京大学,深圳研究生院,深圳,中国)
专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 SSP通过整合光学流和文本提示,提升音频视觉语义分割的精度与效率。