Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
评估检索增强生成与长上下文输入用于电子健康记录临床推理的效果
Skatje Myers, Dmitriy Dligach, Timothy A. Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew Churpek, Anoop Mayampurath, Majid Afshar
机构
*
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Loyola University Chicago(芝加哥洛约拉大学)
;
Boston Children’s Hospital Harvard Medical School(波士顿儿童医院哈佛医学院)
;
University of Colorado-Anschutz(科罗拉多大学安舒茨分校)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
衡量LLM中的推理质量:一个多维行为框架
Ali Şenol, Garima Agrawal, Huan Liu
机构
*
Department of Computer Engineering, Tarsus University(塔鲁斯大学计算机工程系)
;
School of Computing and Augmented Intelligence (SCAI), Arizona State University (ASU)(计算与增强智能学院(SCAI),亚利桑那州立大学(ASU))
;
HumaConn AI Consulting(HumaConn AI咨询)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
BenGER:德国法律中基于归入的法律推理的LLM系统基准测试
Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Koçak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair
机构
*
Technical University of Munich (TUM)(慕尼黑技术大学)
;
Ludwig Maximilian University of Munich (LMU)(慕尼黑路德维希-马克西米利安大学)
;
University of Konstanz(康斯坦茨大学)
;
University of Saarbrücken(萨尔布吕肯大学)
CommentsSpotlight Paper, Proceedings of the Workshop on Structured Data for Health at the 43rd International Conference on Machine Learning (ICML), Seoul, South Korea
机构
*
Department of Computer Science and Software Engineering, Auburn University(奥本大学计算机科学与软件工程系)
;
Department of Information Management, National Central University(国立中央大学信息管理系)
AI总结
提出Age of LLM基准,在13x7网格上让两个LLM进行1v1对战,包含战争迷雾、完全外交和严格JSON格式可靠性维度,通过54场比赛评估15个推理模型,发现核冲刺主导、外交频繁但极少达成、非法动作反映信念追踪,并初步关联可靠性与获胜。
Comments25 pages including appendices, 8 figures, 4 tables; appendices include verbatim system prompt and engine resolution pseudocode. All correlations reported with p-values, 95% bootstrap confidence intervals and Spearman's rho; includes a Steiger test and Bradley-Terry fit
The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act
欧盟法律自动化中的测量差距:欧盟AI法案下教义性法律推理的基准测试
Michèle Finck
机构
*
Chair of Law and Artificial Intelligence and Director, CZS Institute for Artificial Intelligence and Law, University of Tübingen(法律与人工智能教授、人工智能与法律研究所主任,图宾根大学)
CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
CycliST:用于循环状态转换推理的视频语言模型基准
Simon Kohaut, Daniel Ochs, Shun Zhang, Benedict Flade, Julian Eggert, Kristian Kersting, Devendra Singh Dhami
机构
*
Artificial Intelligence and Machine Learning Lab, TU Darmstadt(人工智能与机器学习实验室,图腾斯达特技术大学)
;
Konrad Zuse School of Excellence in Learning and Intelligent Systems (ELIZA)(Konrad Zuse 学校(ELIZA))
;
Honda Research Institute Europe GmbH, Offenbach, Germany(本田欧洲研究院,奥芬巴赫,德国)
;
Uncertainty in Artificial Intelligence Group, TU Eindhoven(人工智能不确定性小组,埃因霍温技术大学)
;
Hessian Center for AI (hessian.AI)(黑森人工智能中心(hessian.AI))
;
Center for Cognitive Science(认知科学中心)
;
German Center for Artificial Intelligence (DFKI)(德国人工智能中心(DFKI))
ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
ClinHallu: 用于诊断医学多模态大语言模型推理中阶段式幻觉的基准
Sicheng Yang, Hangjie Yuan, Wenjun Zhang, Jinwang Wang, Yichen Qian, Weihua Chen, Fan Wang, Lei Zhu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
DAMO Academy, Alibaba Group(阿里巴巴达摩院)
;
Hupan Lab(湖畔实验室)
;
Zhejiang University(浙江大学)
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
LLM能否准确评分医学诊断和临床推理?
Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
机构
*
Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非)
;
Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)
Comments9 pages main text, 31 pages total (including references and appendix). 5 figures, 16 tables. Preprint under review. Code and data will be made available upon publication
机构
*
Department of Computer Science, National Centre for Text Mining, The University of Manchester(计算机科学系,国家文本挖掘中心,曼彻斯特大学)
;
ELLIS Manchester(曼彻斯特ELLIS)
;
School of Computing, Queen’s University, Ontario, Canada(计算学院,加拿大皇后大学)
;
Computer Science, University of Illinois Chicago(计算机科学,伊利诺伊大学芝加哥分校)
;
ELLIS Institute Finland(芬兰ELLIS研究所)
;
University of Turku(图尔库大学)
;
Department of Computer Science and Artificial Intelligence, Umm Al-Qura University, Makkah, Saudi Arabia(计算机科学与人工智能系,乌姆·阿勒·卡拉大学,麦加,沙特阿拉伯)
How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures
语言模型如何失败:承诺性和持续性推理错误的令牌级特征
Tanvi Thoria, Kiana Jafari, Marc R. Schlichting, Mykel J. Kochenderfer
机构
*
Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)
;
Department of Aeronautics and Astronautics, Stanford University(航空航天工程系,斯坦福大学)