机构
*
Shandong University(山东大学)
;
Boston University(波士顿大学)
;
North China Electric Power University(华北电力大学)
;
National University of Singapore(新加坡国立大学)
;
Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University(山东大学-南洋理工大学人工智能联合研究中心(C-FAIR))
;
Rizhao Steel Holding Group Co., Ltd.(日照钢铁控股集团有限公司)
Comments16 pages, 3 figures, 7 tables. Substantially revised: reframed around an identifiability result for the counterbalanced concentration index, with access configuration presented as one generator of the failure rather than a separate confound. Code and data are included as ancillary files
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
PortBench: 一种相关性感知的、全流水线的LLM驱动投资组合管理基准
Yuxuan Zhao, Sijia Chen, Ningxin Su
机构
*
Yantai Research Institute of Harbin Engineering University(哈尔滨工程大学烟台研究院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
Department of Computer and Information Science, University of Pennsylvania(宾夕法尼亚大学计算机与信息科学系)
;
Department of Mathematics, University of Pennsylvania(宾夕法尼亚大学数学系)
Comments16 pages, 9 figures. v2: updated author list (added Yang Liu and Yuxin Li; marked core contributors and project lead) and added a release-date note on the first page
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
MEDLEY-BENCH:在AI元认知中规模购买评估但不控制
Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
机构
*
Department of Clinical Science, Intervention and Technology (CLINTEC), Karolinska Institutet(临床科学、干预与技术部门(CLINTEC),Karolinska研究所)
;
Department of Clinical Physiology, Karolinska University Hospital(临床生理学部门,Karolinska大学医院)
;
Department of Biomedical Engineering and Health Systems, KTH Royal Institute of Technology(生物医学工程与健康系统部门,KTH皇家理工学院)
;
Department of Textile Technology, University of Borås(纺织技术部门,Borås大学)
;
Department of Medical Technologies, Karolinska University Hospital(医学技术部门,Karolinska大学医院)
"LLM Agent Performance" Is Not a Single Evaluation Target
基于LLM的智能体评估统一框架的必要性
Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
机构
*
National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程技术研究中心,武汉大学计算机学院)
;
College of Computing and Data Science, Nanyang Technological University(computing and Data Science学院,南洋理工大学)