Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
见仁见智:在标签噪声下鲁棒的视觉引导跨模态提示学习
机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; PCALab, VCIP, College of Computer Science, Nankai University(南开大学计算机学院PCALab, VCIP)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出VisPrompt框架,通过跨模态注意力机制将视觉语义注入提示表示,提升在标签噪声下的鲁棒性,实验表明其在多个数据集上表现优于现有基线。