CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
CUAAudit: 视觉语言模型作为自主计算机使用代理的元评估
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.AI
AI总结 本文研究了视觉语言模型作为自主计算机使用代理审计员的元评估,发现尽管先进模型在准确性方面表现良好,但在复杂环境中性能下降,突显了现实部署中需考虑评估者可靠性与不确定性的重要性。
Comments This work has been accepted to appear at the HEAL @ CHI 2026 Worshop on Human-centered Evaluation and Auditing of Language Models