Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
实现对话代理的多维评估:一种具有选择性重新评估和模型基准测试的可扩展、受治理的管道
机构 * Lowes(劳氏公司)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
AI总结 研究针对零售对话代理评估难题,提出GenAI Evaluation管道,通过规范化等处理生产日志,能评估多维度指标。选择性重新评估提高效率,支持审计。每日处理约50000条记录,已评估超两百万次交互,取得较好分数和准确率。
Comments 14 pages, 1 figure of design of architecture and 2 tables exploring the results and benchmarking