End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians
面向临床医生的嵌入式电子健康记录(EHR)AI代理的端到端评估与治理
Aaryan Shah, Andrew Hines, Alexia Downs, Denis Bajet, Paulius Mui, Fabiano Araujo, Laura Offutt, Aida Rutledge, Elizabeth Jimenez
机构
*
Canvas Medical
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
XPC (X Primary Care)(XPC(X初级医疗))
;
FCA Consulting(FCA咨询)
;
Department of Pediatrics, University of Nevada(内华达大学儿科系)
How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
AI代理如何花费你的钱?分析和预测代理编码任务中的令牌消耗
Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
机构
*
University of Michigan(密歇根大学)
;
Stanford University(斯坦福大学)
;
All Hands AI
;
Google Deepmind(谷歌DeepMind)
;
Microsoft AI(微软AI)
;
Massachusetts Institute of Technology(麻省理工学院)
Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, Robert Trager
机构
*
Oxford Martin AI Governance Initiative, University of Oxford(牛津大学人工智能治理倡议)
;
MIT Computer Science and Artificial Intelligence Laboratory, MIT(麻省理工学院计算机科学与人工智能实验室)
;
MIT Future Tech(麻省理工学院未来技术)
;
Stanford University(斯坦福大学)
;
Governance and Responsible AI Lab, Purdue University(普渡大学治理与负责任的人工智能实验室)
;
University of Toronto(多伦多大学)
;
Mercatus Center, George Mason University(乔治·马歇尔大学麦卡锡中心)
;
Vilnius University(维尔纽斯大学)
;
Vijil
;
SaferAI
;
AI Standards Lab(人工智能标准实验室)
;
The Future Society(未来社会)
;
Concordia AI(康科德人工智能)
;
Pivotal Research
;
Center for AI Risk Management & Alignment(人工智能风险管理和对齐中心)
;
UC Berkeley Center for Long-Term Cybersecurity(伯克利大学长期网络安全中心)
;
Independent(独立)
机构
*
Stanford University(斯坦福大学)
;
Shenzhen University(深圳大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong(香港城市大学)
;
Renmin University of China(中国人民大学)
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
少做决定,多进行沟通:关于医学领域端到端事实核查构建效度的探讨
Sebastian Joseph, Lily Chen, Barry Wei, Michael Mackert, Iain J. Marshall, Paul Pu Liang, Ramez Kouzy, Byron C. Wallace, Junyi Jessy Li
机构
*
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Stanford University(斯坦福大学)
;
Indiana University School of Medicine(印第安纳大学医学院)
;
King’s College London(伦敦国王学院)
;
Massachusetts Institute of Technology(麻省理工学院)
;
The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
;
Northeastern University(东北大学)
Automated Adversarial Collaboration for Advancing Theory Building in the Cognitive Sciences
自动化对抗协作促进认知科学理论构建
Suyog Chandramouli, George Kachergis, Akshay Jagadish
机构
*
Department of Psychology, Princeton University(普林斯顿大学心理学系)
;
Department of Psychology, Stanford University(斯坦福大学心理学系)
;
Princeton AI Lab, Princeton University(普林斯顿大学人工智能实验室)
Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis
增强财务报告问答:一种带有重排序分析的检索增强生成系统
Zhiyuan Cheng, Longying Lai, Yue Liu, Kai Cheng, Xiaoxi Qi
机构
*
School of Engineering Stanford University Stanford, CA, USA
;
Simon Business School University of Rochester Rochester, NY, USA
;
Accounting \& Information Systems Rutgers University Newark, NJ, USA
;
Institute for Social
;
Economic Research
;
Policy Columbia University New York, NY, USA
;
Department of Economics Northeastern University Boston, MA, USA
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
Gesture2Music: 一种低延迟的实时框架,用于连续的手势驱动音乐生成
Rathinaraja Jeyaraj, Barathi Subramanian, Kapilya Gangadharan, Anand Paul
机构
*
Stanford University(斯坦福大学)
;
Saveetha Institute of Medical and Technical Sciences(Saveetha医学与技术科学学院)
;
LSU Health Sciences Center New Orleans(路易斯安那州立大学健康科学中心新奥尔良分校)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
针对具体病例的评估量表:方法、验证及LLM与医生一致性的823次病例研究
Aaryan Shah, Andrew Hines, Alexia Downs, Denis Bajet, Paulius Mui, Fabiano Araujo, Laura Offutt, Aida Rutledge, Elizabeth Jimenez
机构
*
Canvas Medical
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
XPC (X Primary Care)(XPC(X初级保健))
;
FCA Consulting(FCA咨询)
;
Department of Pediatrics, University of Nevada, Reno, NV, USA(内华达大学拉斯维加斯分校儿科系)
Cooptimizing Safety and Performance Using Safety Value-Constrained Model Predictive Control
通过安全价值约束模型预测控制优化安全与性能
Hao Wang, Nam Nguyen, Armand Jordana, Ludovic Righetti, Somil Bansal
机构
*
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California(南加州大学明希德电气与计算机工程系)
;
Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空与航天系)
;
Electrical and Computer Engineering Department, New York University(纽约大学电气与计算机工程系)
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
FunRec:从第一人称交互视频重建功能3D场景
Alexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario, Boyang Sun, Rishabh Dabral, Leonidas Guibas, Christian Theobalt, Marc Pollefeys, Francis Engelmann, Daniel Barath
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学))
;
Stanford University(斯坦福大学)
;
Microsoft(微软)
;
USI Lugano(USI卢加诺)
Using Language Models as Closed-Loop High-Level Planners for Robotics Applications: A Brief Overview and Benchmarks
利用语言模型作为闭环高层规划器用于机器人应用:简要概述与基准测试
Hao Wang, Sathwik Karnik, Bea Lim, Somil Bansal
机构
*
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California(南加州大学明希德电气与计算机工程系)
;
Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空航天工程系)
;
Department of Mechanical Engineering, Stanford University(斯坦福大学机械工程系)