Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence
跨领域基准测试揭示协调AI代理在部分证据下提升科学推断何时有效
Fiona Y. Wong, Markus J. Buehler
机构
*
Laboratory for Atomistic and Molecular Mechanics (LAMM)(原子分子力学实验室)
;
Department of Biological Engineering(生物工程系)
;
Department of Mechanical Engineering(机械工程系)
;
Department of Civil and Environmental Engineering(土木与环境工程系)
;
Center for Computational Science and Engineering, Schwarzman College of Computing(计算科学与工程中心)
;
Massachusetts Institute of Technology(麻省理工学院)
机构
*
New York University Abu Dhabi(纽约大学阿布扎克校区)
;
Nanyang Technological University(南洋理工大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
The Hong Kong Polytechnic University(香港理工大学)
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
STRUCTSENSE:一种任务无关的代理框架,用于结构化信息提取,具有人机协同评估和基准测试
Tek Raj Chhetri, Yibei Chen, Puja Trivedi, Dorota Jarecka, Saif Haobsh, Patrick Ray, Lydia Ng, Satrajit S. Ghosh
机构
*
McGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, MA, USA(麦戈文脑科学研究所,麻省理工学院,马萨诸塞州剑桥市)
;
Fylo Labs Inc., New York, NY, USA(Fylo实验室公司,纽约州纽约市)
;
Allen Institute for Brain Science, Seattle, WA, USA(艾伦脑科学研究所,华盛顿州西雅图市)
AOP-Wiki EMOD 3.0: Data Model Expansions and Content Evaluation Framework for Using Agentic AI to Improve Integration between AOPs and New Approach Methodologies (NAMs)
Virginia K. Hench, J. Harry Caufield, Sierra A. T. Moxon, Jason M. O'Brien, Stephen W. Edwards
机构
*
Open BioData Modeling(开放生物数据建模)
;
Environmental Genomics and Systems Biology(环境基因组学与系统生物学)
;
Lawrence Berkeley National Laboratory(伯克利国家实验室)
;
National Wildlife Research Centre(国家野生动物研究中心)
;
UL Research Institutes - Chemical Insights(UL研究机构-化学洞察)
机构
*
New York University Abu Dhabi(纽约大学阿布扎克校区)
;
Nanyang Technological University(南洋理工大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Harvard University(哈佛大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing University of Technology(北京理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
Comments9 pages main text plus appendices, 8 figures. Dataset and benchmark paper. ChronoMedKG released under CC BY 4.0 and ChronoTQA/code under MIT (Zenodo: 10.5281/zenodo.19697542). Under review
CommentsMingyuan served as the project lead. Banghao, Yining, and Mingyuan contributed equally to this work, with more junior authors listed before senior authors. All data and code releases are maintained by the corresponding authors at UIUC and are not affiliated with Meta
Comments18 pages, 5 figures, 1 appendix. Open-source artifact at github.com/mrdaemoni/myalicia (MIT). Preregistered multi-participant replication study planned on OSF. Companion essay "The Humorphic Partnership" at myalicia.com. Design philosophy at humorphism.com
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
在不欺骗自己的情况下衡量安全:为什么基准测试智能体是困难的
Sahar Abdelnabi, Chris Hicks, Konrad Rieck, Ahmad-Reza Sadeghi
机构
*
ELLIS Institute Tübingen & MPI-IS & Tübingen AI Center(图宾根ELLIS研究所及MPI-IS与图宾根人工智能中心)
;
The Alan Turing Institute(艾伦·图灵研究所)
;
BIFOLD & Technische Universität Berlin(BIFOLD与柏林技术大学)
;
Technische Universität Darmstadt(达姆施塔特技术大学)
Houxuan Zhou, Sriram Prasad, Chenghao Huang, Jiajie Feng, Hao Wang
机构
*
Department of Data Science and AI, Faculty of IT, Monash University, Australia(数据科学与人工智能系,IT学院,墨尔本大学,澳大利亚)
;
School of Electrical Engineering and Computer Science, University of Queensland, Australia(电气工程与计算机科学学院,昆士兰大学,澳大利亚)
;
Monash Energy Institute, Monash University, Australia(墨尔本能源研究所,墨尔本大学,澳大利亚)