VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild
VLMGuard:从野生未标记视觉语言提示中引导恶意提示检测器
机构 * College of Computing and Data Science(计算与数据科学学院) ; Nanyang Technological University(南洋理工大学) ; School of Physical and Mathematical Sciences(物理与数学科学学院) ; Microsoft Corp.(微软公司) ; Department of Computer Sciences(计算机科学系) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
AI总结 研究针对视觉语言模型易受恶意输入影响的问题,提出VLMGuard框架,利用野生未标记用户提示,通过自动恶意估计分数区分良性和恶意样本,训练二进制提示分类器,无需额外人工标注,效果优于现有方法。
Comments Accepted to Transactions on Machine Learning Research (07/2026)