Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
机构 * Hefei University of Technology(合肥工业大学) ; Chinese Academy of Sciences(中国科学院) ; Beihang University(北航) ; Sangfor Technologies(深信服技术)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
Comments Accepted by CVPR 2025; Code: https://github.com/spyflying/VCT_AVS; Models: https://huggingface.co/nowherespyfly/VCT_AVS