Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
大型语言模型(LLMs)的财务推理是否可信?针对长期财务报表的现实测试
专题命中 安全评测 :alignment(abstract);分类 cs.CL
AI总结 该研究针对LLMs的财务推理可信度,构建含对抗陷阱的FinIndices基准测试,发现其存在知识与结构瓶颈,监督微调可部分恢复结构化逻辑。
Comments The FinIndices dataset is publicly available at https://huggingface.co/datasets/User158072/Finindice