FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
FrontierFinance: 一个长期计算机使用的真实金融任务基准
Michael Krumdick, Varshini Reddy, Shivani Chaudhary, William Day, Maarij Ahmed, Hayan Haqqi, Muhammad Ahsen Fahim, Hanzallah Amjad, Ahmad Orakzai, Aqsa Gul, Chris Tanner
Blind-Spot Mass: A Good-Turing Framework for Quantifying Deployment Coverage Risk in Machine Learning Systems
盲区质量:一种用于量化机器学习系统部署覆盖风险的Good-Turing框架
Biplab Pal, Santanu Bhattacharya, Madanjit Singh
机构
*
University of Maryland, Baltimore County (UMBC)(马里兰大学巴尔的摩县分校)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Ambient Scientific Inc(Ambient Scientific公司)
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
TeamPath: 构建多模态病理专家的推理AI助手
Tianyu Liu, Weihao Xuan, Hao Wu, Peter Humphrey, Marcello DiStasio, Mohamed Kahila, Alfonso Garcia Tan, Heli Qi, Rui Yang, Simeng Han, Tinglin Huang, Fang Wu, Chen Liu, Qingyu Chen, Nan Liu, Irene Li, Hua Xu, Hongyu Zhao
机构
*
Interdepartmental Program in Computational Biology and Biomedical Informatics, Yale University(耶鲁大学计算生物学与生物医学信息学跨学科项目)
;
Department of Biostatistics, Yale University(耶鲁大学生物统计学系)
;
Broad Institute of MIT and Harvard(博德研究所)
;
Department of Complexity Science and Engineering, The University of Tokyo(东京大学复杂科学与工程系)
;
Center for Advanced Intelligence Project, RIKEN(理化学研究所先进智能项目中心)
;
Department of Pathology, Yale University(耶鲁大学病理学系)
;
Department of Anatomical Pathology, Singapore General Hospital(新加坡中央医院解剖病理学系)
;
Center for Biomedical Data Science, Duke–NUS Medical School, Singapore, Singapore(杜克-新加坡国立大学医学院生物医学数据科学中心)
;
Department of Computer Science, Yale University(耶鲁大学计算机科学系)
;
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
机构
*
Yale University(耶鲁大学)
;
Broad Institute of MIT and Harvard(麻省理工学院-哈佛大学博德研究所)
;
Google DeepMind(谷歌DeepMind)
;
Stanford University(斯坦福大学)
;
Genentech(基因泰克)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Cornell University(康奈尔大学)
;
Harvard University(哈佛大学)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
对冲与非肯定:量化大语言模型在人权问题上的对齐
Rafiya Javed, Cassandra Parent, Jackie Kay, David Yanni, Abdullah Zaini, Anushe Sheikh, Maribeth Rauh, Walter Gerych, Ramona Comanescu, Iason Gabriel, Marzyeh Ghassemi, Laura Weidinger
机构
*
Google Deepmind(谷歌DeepMind)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Independent Researcher(独立研究员)
;
Google(谷歌)
;
AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)