DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
DFAH-Bench:金融决策中可观测智能体不稳定性的基准测试
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究金融决策中智能体行为稳定性,引入DFAH-Bench基准,通过多渠道衡量。发现仅结果一致性不能完全反映稳定性,前沿模型存在决策与工具路径一致性差距,识别出三种行为模式,并开源相关代码和数据。
Comments 15 pages, 3 figures. Code, sanitized replay logs, one-command reproduction (make reproduce-paper), and an interactive results explorer: https://github.com/ibm-client-engineering/output-drift-financial-llms