AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
AutoGUI-v2:一个全面的多模态GUI功能理解基准
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; New Laboratory of Pattern Recognition(模式识别新实验室) ; State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) ; Hong Kong Institute of Science & Innovation(香港科学创新研究院) ; PolyU ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
AI总结 AutoGUI-v2通过多平台截图递归解析生成多样化任务,评估深度GUI功能理解和交互预测能力,揭示VLMs在功能接地与描述上的差异及复杂交互逻辑的挑战。
Comments Technical Report