Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness
Agent-UCT:应用于树的上置信界,用于具有成本意识的智能工作流优化
Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
机构
*
The University of Hong Kong(香港大学)
;
School of Artificial Intelligence, Jiangxi Science and Technology Normal University(江西科技师范大学人工智能学院)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
University of Science and Technology of China(中国科学技术大学)
;
New York University(纽约大学)
CommentsThe authors have withdrawn this article because the current version is still undergoing substantial revision. Several components of the data synthesis framework, consistency-filtering procedure, evaluation protocol, and experimental analysis are being refined and expanded. As a result, the current manuscript should not be considered a complete or final representation of the work
Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles
告知、指导、共情、倾听:审计LLM护理支持角色
Drishti Goel, Agam Goyal, Veda Duddu, Olivia Pal, Jeongah Lee, Qiuyue Joy Zhong, Violeta J. Rodriguez, Daniel S. Brown, Dong Whi Yoo, Ravi Karkar, Koustuv Saha
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
;
OSF HealthCare(OSF医疗集团)
;
Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
首先,不伤害:迈向临床安全的大语言模型
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
机构
*
Harvard Combined Dermatology Program(哈佛联合皮肤科项目)
;
Department of Dermatology, Mass General Brigham(麻省总医院皮肤科)
;
Harvard Medical School(哈佛医学院)
;
Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心)
;
Stanford University(斯坦福大学)
;
Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科)
;
Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科)
;
Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯)
;
Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科)
;
Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科)
;
Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科)
;
Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科)
;
Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科)
;
Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科)
;
Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科)
;
Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心)
;
Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所)
;
Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
Centre for Perceptual and Interactive Intelligence (CPII)(感知与交互智能中心)
;
OceanBase
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
需要多少次迭代才能突破限制?多轮LLM评估中的动态预算分配
Shai Feldman, Yaniv Romano
机构
*
Department of Computer Science(计算机科学系)
;
Technion, Israel(技术ion, 以色列)
;
Departments of Electrical and Computer Engineering and of Computer Science(电气与计算机工程系和计算机科学系)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
MEDIC:对LLM在临床应用中的安全性和实用性领先指标的综合评估
Praveenkumar Kanithi, Clément Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan
MedBeads: An AI-Native Clinical Context Graph Built from Immutable Beads and Reconstructable Clinical Links
MedBeads:面向可信医疗AI的智能体原生不可变数据基底
Takahito Nakajima
机构
*
Diagnostic Imaging and Interventional Radiology, Institute of Medicine, University of Tsukuba(东京大学医学研究院诊断影像与介入放射学部)
;
Center for Cyber Medicine Research, University of Tsukuba(东京大学计算机医学研究中心)
Comments21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval