Comments46 pages including appendices (two-column preprint format). Under review at JAMIA. Code, frozen evaluator, and benchmark released at https://huggingface.co/datasets/alexstinard/epikg-clinicalbench. ClinicalBench v2 is a 400-question MIMIC-IV stress test for assertion-aware retrieval
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
ExploitGym:AI代理能否将安全漏洞转化为实际攻击?
Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song
机构
*
UC Berkeley(加州大学伯克利分校)
;
Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所)
;
UC Santa Barbara(加州大学圣巴巴拉分校)
;
Arizona State University(亚利桑那州立大学)
;
Anthropic(Anthropic公司)
;
OpenAI
;
Google(谷歌)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Stargazer:一种可扩展的AI代理在天体物理约束下的模型拟合基准环境
Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin, Kristen Menou
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
RW-Post:可审计的证据导向多模态事实核查
Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli
机构
*
School of Computing (SoC), National University of Singapore (NUS)(新加坡国立大学计算机学院(SoC))
;
National University of Singapore (NUS)(新加坡国立大学)
;
Department of Electrical and Computer Engineering (ECE), National University of Singapore (NUS)(新加坡国立大学电子与计算机工程系(ECE))
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
不要Pass@k:大规模语言模型评估的贝叶斯框架
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin Chaudhary
机构
*
Department of Computer and Data Sciences, Case Western Reserve University(计算机与数据科学系,凯斯西储大学)
;
Department of Physics, Case Western Reserve University(物理系,凯斯西储大学)
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳瓦分校计算机科学与工程系)
;
School of Computer Engineering, KIIT Deemed to be University, Bhubaneswar, India(比哈尔邦布尔萨大学计算机工程学院)
Speech-based Psychological Crisis Assessment using LLMs
基于语音的心理危机评估使用大语言模型
Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang
机构
*
Tsinghua University(清华大学)
;
Peking University Huilongguan Clinical Medical School(北京大学回龙guan临床医学院)
;
WHO Collaborating Centre for Research and Training in Suicide Prevention(世界卫生组织自杀预防研究与培训协作中心)
NaiAD: Initiate Data-Driven Research for LLM Advertising
NaiAD:启动基于数据的研究以进行大语言模型广告
Yihang Zhang, Zimeng Huang, Ren Zhai, Yipeng Kang, Tonghan Wang
机构
*
Tsinghua University(清华大学)
;
College of AI(人工智能学院)
;
Department of Literature, Arts and Communication(文学、艺术与传播系)
;
Anhui International Studies University(安徽国际关系大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
Towards Conversational Medical AI with Eyes, Ears and a Voice
面向有眼睛、耳朵和声音的对话式医疗AI
Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院)
;
Stanford University(斯坦福大学)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
生成无泄漏的基准以评估鲁棒的RAG
Jiayi Liu, Jiaxing Zhang, Bowen Jin, Jennifer Neville
机构
*
Department of Computer Science, Purdue University(普渡大学计算机科学系)
;
New Jersey Institute of Technology(新泽西理工学院)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Microsoft Research(微软研究院)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
基准是否低估了大语言模型的性能?通过大语言模型优先的人类仲裁评估来评估幻觉检测
I. F. Atasoy, B. Mutlu, E. A. Sezer, A. Wahdan
机构
*
Department of Computer Engineering, Hacettepe University(哈切塔佩大学计算机工程系)
;
Department of Computer Engineering, Ankara University(安卡拉大学计算机工程系)
;
Zephlen AI and Information Technologies Inc.(泽夫伦人工智能与信息技术公司)
Commentshis is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: https://doi.org/10.1109/SUMMA68668.2025.11302248
Journal ref2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA), Lipetsk, Russian Federation, 2025, pp. 799-804
机构
*
National Chengchi University(国立中正大学)
;
Georgetown University(乔治城大学)
;
University of Michigan(密歇根大学)
;
Stevens Institute of Technology(史蒂文斯理工学院)
;
National Taiwan University of Science and Technology(台湾科技大学)
;
Far Eastern Memorial Hospital(东方纪念医院)
Evaluating Large Language Models in Scientific Discovery
评估大型语言模型在科学发现中的表现
Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo, Yuanhao Qu, Yi Qu, Yinkai Wang, Haorui Wang, Jeff Guo, Jingru Gan, Parshin Shojaee, Di Luo, Andres M Bran, Gen Li, Qiyuan Zhao, Shao-Xiong Lennon Luo, Yuxuan Zhang, Xiang Zou, Wanru Zhao, Yifan F. Zhang, Wucheng Zhang, Shunan Zheng, Saiyang Zhang, Sartaaj Takrim Khan, Mahyar Rajabi-Kochi, Samantha Paradi-Maropakis, Tony Baltoiu, Fengyu Xie, Tianyang Chen, Kexin Huang, Weiliang Luo, Meijing Fang, Xin Yang, Lixue Cheng, Jiajun He, Soha Hassoun, Xiangliang Zhang, Wei Wang, Chandan K. Reddy, Chao Zhang, Zhiling Zheng, Mengdi Wang, Le Cong, Carla P. Gomes, Chang-Yu Hsieh, Aditya Nandy, Philippe Schwaller, Heather J. Kulik, Haojun Jia, Huan Sun, Seyed Mohamad Moosavi, Chenru Duan
机构
*
Deep Principle(深原则)
;
Department of Computer Science, Cornell University(计算机科学系,康奈尔大学)
;
Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学)
;
Department of Chemical Engineering & Applied Chemistry, University of Toronto(化学工程与应用化学系,多伦多大学)
;
Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)
;
QuEra Computing Inc.(QuEra计算公司)
;
Department of Pathology, Department of Genetics, Cancer Biology Program, Stanford University School of Medicine(病理学系、遗传学系、癌症生物学项目,斯坦福大学医学院)
;
Harvard Law School(哈佛法学院)
;
Department of Computer Science, Tufts University(计算机科学系,塔夫茨大学)
;
School of Computational Science and Engineering, Georgia Institute of Technology(计算科学与工程学院,佐治亚理工学院)
;
Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)
;
Department of Computer Science, Virginia Tech(计算机科学系,弗吉尼亚理工大学)
;
Department of Physics, Tsinghua University(物理系,清华大学)
;
Institute for Advanced Study, Tsinghua University(清华大学高级研究所)
;
Laboratory of Artificial Chemical Intelligence, Ecole Polytechnique Federale de Lausanne(人工化学智能实验室,瑞士联邦理工学院)