CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
CUJBench:跨模态故障诊断的LLM-代理基准测试
专题命中 GUI与网页智能体 :agent(title,abstract);multi-agent(abstract);分类 cs.SE
AI总结 CUJBench是首个结合浏览器可见故障证据与后端可观测性的跨模态故障诊断基准测试,通过LLM辅助生成管道降低标注成本,评估六种前沿模型发现跨模态综合是主要瓶颈。
Comments 10 pages, 1 figure; updated source code url