Comments20 pages, 3 figures, 5 tables. Accepted at the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). To appear in LIPIcs Vol. 394. v2: corrected the bibliographic record of one reference (preprint, not a journal article) and added the related-version link to the published LIPIcs article
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
TimeSage-EV:面向动态环境下智能体时间序列分析的实时基准测试集
Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo, Qingsong Wen, Ming Jin, Joaquin Vanschoren
机构
*
Eindhoven University of Technology(埃因霍温理工大学)
;
University of Oxford(牛津大学)
;
VulpiVox Intelligence(VulpiVox智能公司)
;
Squirrel Ai Learning(Squirrel AI学习公司)
;
Griffith University(格里菲斯大学)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
AnchorBench:针对大语言模型中锚定效应的多路径基准测试
Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin
机构
*
Saarland University(萨尔大学)
;
Hamburg University of Technology(汉堡工业大学)
;
Helmholtz-Zentrum Hereon(亥姆霍兹中心赫伦)
;
German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
专题命中
评测与基准
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
手术内窥镜视频的时间视觉语言模型的鲁棒性研究
Darakshan Rashid, Raza Imam, Ufaq Khan, Muhammad Bilal, Shazad Ashraf, Dwarikanath Mahapatra, Mohammad Yaqub, Muhammad Haris Khan, Imran Razzak, Brejesh Lall, Lena Maier-Hein, Yutong Xie
机构
*
Indian Institute of Technology Delhi(印度理工学院德里分校)
;
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Birmingham City University(伯明翰城市大学)
;
University Hospitals Birmingham(伯明翰大学医院)
;
Khalifa University(哈利法大学)
;
German Cancer Research Center (DKFZ)(德国癌症研究中心)