SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
SciVisAgentBench:用于评估科学数据分析和可视化代理的基准测试
机构 * University of Notre Dame(圣母大学) ; Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) ; University of Utah(犹他大学) ; University of Nebraska–Lincoln(内布拉斯加大学林肯分校) ; The Ohio State University(俄亥俄州立大学) ; Argonne National Laboratory(阿贡国家实验室) ; Anthropic PBC
专题命中 代码评测 :coding agent(abstract);分类 cs.AI
AI总结 本文提出SciVisAgentBench,一个用于评估科学数据可视化代理的基准测试,涵盖四个维度,包含108个案例,结合LLM和确定性评估器进行多模态评估,揭示代理能力差距。
Comments IEEE Transactions on Visualization and Computer Graphics (IEEE VIS '26)