Journal refProceedings of the 38th International Conference on Neural Information Processing Systems (NIPS '24), Vol. 37 (2024), Article 3551, 111832-111862
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Shu Liu, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Jasjeet Sekhon, Jacob Steinhardt, Antony Kellermann, Sarah Schwettmann, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang
机构
*
UIUC(伊利诺伊大学)
;
Stanford University(斯坦福大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
Yale University(耶鲁大学)
;
Princeton University(普林斯顿大学)
;
MIT(麻省理工学院)
;
Transluce
;
ML Commons
;
Amazon(亚马逊)
;
UK AI Safety Institute(英国人工智能安全研究所)
;
University of Oxford(牛津大学)
机构
*
Business School, Hunan University(湖南大学商学院)
;
Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)
;
School of Business, Hunan University(湖南大学商学院)
;
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(香港科技大学(香港)计算机科学与工程学院)
AI agents may be worth the hype but not the resources (yet): An initial exploration of machine translation quality and costs in three language pairs in the legal and news domains