Offline Two-Player Zero-Sum Markov Games with KL Regularization
离线双玩家零和马尔可夫游戏中的KL正则化
Claire Chen, Yuheng Zhang, Xinyu Liu, Zixuan Xie, Shuze Daniel Liu, Nan Jiang
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
California Institute of Technology(加州理工学院)
;
University of Virginia(弗吉尼亚大学)
;
Purdue University(Purdue 大学)
;
Massachusetts Institute of Technology(麻省理工学院)
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Hong Kong University of Science and Technology(香港科学与技术大学)
;
Texas A&M University(德克萨斯农工大学)
RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records
RDMA: 低成本的代理驱动罕见病挖掘从电子健康记录
John Wu, Adam Cross, Jimeng Sun
机构
*
Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
Department of Pediatrics, University of Illinois College of Medicine Peoria(伊利诺伊大学皮奥里亚医学院儿科系)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
MedHopQA: 一种以疾病为中心的多跳推理基准和评估框架,用于基于大语言模型的生物医学问答
Rezarta Islamaj, Robert Leaman, Joey Chan, Nicholas Wan, Qiao Jin, Natalie Xie, John Wilbur, Shubo Tian, Lana Yeganova, Po-Ting Lai, Chih-Hsuan Wei, Yifan Yang, Yao Ge, Qingqing Zhu, Zhizheng Wang, Zhiyong Lu
机构
*
National Library of Medicine, Division of Intramural Research(国家医学图书馆,院内研究部)
;
University of Illinois at Urbana-Champaign, Department of Computer Science(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
University of Michigan Medical School(密歇根大学医学院)
Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
BioCreative IX MedHopQA 轨道概述:轨道描述、参与及多跳医学问答系统评估
Rezarta Islamaj, Joey Chan, Robert Leaman, Jongmyung Jung, Hyeongsoon Hwang, Quoc-An Nguyen, Hoang-Quynh Le, Harikrishnan Gurushankar Saisudha, Ganesh Chandrasekar, Rustam R. Taktashov, Nadezhda Yu. Bizyukova, Sofia I. R. Conceição, Paulo R. C. Lopes, Reem Abdel Salam, Mary Adewunmi, Zhiyong Lu
机构
*
National Library of Medicine (NLM), National Institutes of Health (NIH)(美国国家医学图书馆(NLM)、国家卫生研究院(NIH))
;
University of Illinois at Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Korea University(韩国大学)
;
VNU University of Engineering and Technology, Hanoi, Vietnam(越南河内工程大学)
;
Concordia University, Montreal, QC, CA(蒙特利尔大学)
;
Institute of Biomedical Chemistry (IBMC), 10 bld. 8, Pogodinskaya str., 119121 Moscow, Russia(俄罗斯生物医学化学研究所(IBMC))
;
LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisbon, Portugal(葡萄牙里斯本大学 LASIGE 实验室)
;
Faculty of Engineering, Computer Engineering Department Cairo University(埃及开罗大学工程学院)
;
Menzies School of Health Research, Charles Darwin University, NT, Australia(澳大利亚查尔斯达尔文大学梅恩兹健康研究中心)
;
CaresAI, Australia(澳大利亚 CaresAI)
AI总结
本文介绍了BioCreative IX MedHopQA共享任务,旨在评估大型语言模型在多跳推理中的表现,通过构建1000个挑战性问题,展示了检索增强生成策略的重要性,并提供了公开数据集和评估结果。
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
一念之差:多轮对话中针对隐藏恶意意图的响应感知防御
Xinjie Shen, Rongzhe Wei, Peizhi Niu, Haoyu Wang, Ruihan Wu, Eli Chien, Bo Li, Pin-Yu Chen, Pan Li
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
UCSD(加州大学圣地亚哥分校)
;
National Taiwan University(台湾大学)
;
IBM Research(IBM研究院)
;
Virtue AI
Acceleration of horizontal numerical advection for atmospheric modeling through surrogate modeling with temporal coarse-graining
通过时间粗化与代理建模加速大气模型中的水平数值输运
Manho Park, Christopher V. Rackauckas, Christopher W. Tessum
机构
*
The Grainger College of Engineering, Department of Civil and Environmental Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格拉inger工程学院,土木与环境工程系)
;
Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA(麻省理工学院计算机科学与人工智能实验室,剑桥马萨诸塞州,美国)
机构
*
Department of Information Science, University of North Texas(北卡罗来纳州立大学信息科学系)
;
Department of Computer Science, North Carolina State University(北卡罗来纳州立大学计算机科学系)
;
Department of Data Science, University of North Texas(得克萨斯大学数据科学系)
;
School of Information Sciences, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校信息科学学院)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
RubricEM: 通过规则引导的策略分解实现超越可验证奖励的元强化学习
Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Google Cloud AI Research(谷歌云人工智能研究)