PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
PreAct-Bench:大语言模型中的预测性监控基准
Hainiu Xu, Italo Luis da Silva, Jiangnan Ye, Yuhao Wang, Wei Liu, Linyi Yang, Jonathan Richard Schwarz, Nicola Paoletti, Yulan He, Hanqi Yan
机构
*
King’s College London(伦敦国王学院)
;
National University of Singapore(新加坡国立大学)
;
Southern University of Science and Technology(南方科技大学)
;
Thomson Reuters Foundational Research(汤姆森路透基础研究)
;
Imperial College London(伦敦帝国学院)
;
The Alan Turing Institute(艾伦·图灵研究所)
机构
*
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)
;
iFLYTEK AI Research (Central China), iFLYTEK Co., Ltd(iFLYTEK中央中国AI研究院,iFLYTEK公司)
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
大语言模型能否以具有法律意义的方式进行推理?针对欧洲人权法院案件的小规模研究
Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
机构
*
University of Copenhagen(哥本哈根大学)
;
Faculty of Law, University of Copenhagen(哥本哈根大学法学院)
;
Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
School of Data Science, Fudan University(复旦大学数据科学学院)
;
Ant Group(蚂蚁集团)
;
School of Information and School of Smart Governance, Renmin University of China(中国人民大学信息学院与智慧治理学院)