MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
MAG:用于多模态动作与引导生成的网络智能体基准测试与工具包
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Monash University(莫纳什大学) ; Pusan National University(釜山国立大学) ; Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.CL
AI总结 介绍MAG这一网络智能体基准测试,统一任务执行与引导写作,有基于截图的定位方案和完整工具包。用其评估模型并详细分析,还设计GRPO训练方法,提升智能体成功率与引导质量,指出当前模型任务完成率低,为后续研究提供方向。
Comments 8 pages main text, 21 pages total including appendices; 11 figures, 7 tables, 2 algorithms. Benchmark, harness, and model checkpoints to be released