Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
基准测试还不够:RAMP——生产系统中代理模型的运行时评估
机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(中山大学计算机科学与工程学院,广州,中国)
专题命中 工作流自动化 :agentic(title,abstract);workflow(abstract);分类 cs.AI、cs.SE
AI总结 针对现有基准测试无法反映真实生产环境动态复杂性的问题,提出RAMP框架,通过统一运行时评估架构、编译器构建工作负载和多维效用指标,揭示模型在长序列工作流中的性能退化与资源效率差异。
Comments 16 pages, 8 figures. Project homepage: http://ramp.yatcc-ai.com/