TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
TS-Haystack:一种用于长上下文时间序列推理的多任务检索基准
机构 * Agentic Systems Lab, ETH Zurich(1 非常规系统实验室,苏黎世联邦理工学院) ; Stanford University(2 斯坦福大学) ; Traffic Engineering Group, Institute for Transport Planning and Systems, ETH Zurich(3 交通工程组,交通规划与系统研究所,苏黎世联邦理工学院) ; University of Illinois Urbana-Champaign(4 印第安纳大学厄巴纳-香槟分校) ; Google(5 谷歌) ; Centre for Digital Health Interventions, ETH Zurich(6 数字健康干预中心,苏黎世联邦理工学院) ; Centre for Digital Health Interventions, University of St. Gallen(7 数字健康干预中心,圣加尔登大学)
AI总结 本文提出TS-Haystack基准,评估长上下文时间序列推理能力,发现现有TSLMs存在长上下文退化问题,代理检索框架在9/10任务中表现优异。
Comments Workshop version of this paper published at ICLR TSALM 2026. Benchmark generation code and datasets: https://github.com/AI-X-Labs/TS-Haystack