Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
长视野终端基准测试:使用基于密集奖励的评分方式测试智能体在长视野终端任务中的极限
机构 * Tencent(腾讯) ; University of Maryland, College Park(马里兰大学帕克分校) ; University of Georgia(佐治亚大学) ; University of Minnesota, Twin Cities(明尼苏达大学双城分校) ; Indiana University(印第安纳大学) ; Lehigh University(里海大学) ; National University of Singapore(新加坡国立大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态评测 :multimodal(abstract);分类 cs.AI
AI总结 研究针对现有终端基准测试局限,引入长视野终端基准测试(Long-Horizon-Terminal-Bench),含46个长视野任务。通过分解为分级子任务提供密集中间奖励,评估15个前沿模型,揭示改进空间,分析失败模式,发布该基准测试助力长视野终端智能体发展。
Comments 17 pages