EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
EnvSimBench:一个评估和改进基于LLM的环境模拟的基准
Yi Liu, TingFeng Hui, Wei Zhang, Li Sun, Ningxin Su, Jian Wang, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Chongqing University(重庆大学)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
工具选择:测量LLM代理追求工具性行为的倾向
Jonas Wiedermann-Möller, Leonard Dung, Maksym Andriushchenko
机构
*
Universität Bielefeld(比勒菲尔德大学)
;
Ruhr-Universität Bochum(波鸿鲁尔大学)
;
ELLIS Institute(ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Tübingen AI Center(图宾根人工智能中心)
CommentsWe have further refined the benchmark construction and reference verification pipeline to improve clarity and consistency. The revised version includes updated results and additional details to better align the evaluation with the intended setup. These changes provide a more precise presentation of the experimental findings, with conclusions and contributions remaining unchanged
机构
*
East China Normal University(华东师范大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Shanghai Innovation Institute(上海创新研究院)
Aishik Mandal, Hiba Arnaout, Clarissa W. Ong, Juliet Bockhorst, Kate Sheehan, Rachael Moldow, Tanmoy Chakraborty, Iryna Gurevych
机构
*
UKP Lab, Department of Computer Science and Hessian Center for AI (hessian.AI), Technische Universität Darmstadt(德国达姆施塔特技术大学UKP实验室、计算机科学系和海森人工智能中心(hessian.AI))
;
Zuse School ELIZA(Zuse学院ELIZA)
;
National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE)
;
Indian Institute of Technology Delhi(印度德里理工学院)
;
Yardi School of Artificial Intelligence(Yardi人工智能学院)
;
University of Louisville(路易斯维尔大学)
;
University of Toledo(托莱多大学)
Toward Safe Autonomous Robotic Endovascular Interventions using World Models
迈向安全的自主机器人内血管干预的世模型
Harry Robertshaw, Nikola Fischer, Han-Ru Wu, Andrea Walker Perez, Weiyuan Deng, Benjamin Jackson, Christos Bergeles, Alejandro Granados, Thomas C Booth
机构
*
Surgical & Interventional Engineering, School of Biomedical Engineering & Imaging Sciences, King’s College London(外科与介入工程系,生物医学工程与成像科学学院,伦敦国王学院)
;
Department of Radiology, National Taiwan University Hospital(放射科,台湾大学医院)
;
Department of Neuroradiology, King’s College Hospital(神经放射科,伦敦国王医院)
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
大语言模型代理中的不确定性量化:基础、新兴挑战与机遇
Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Xuefeng Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li
机构
*
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of Southern California(南加州大学)
;
University of California, Berkeley(加州大学伯克利分校)
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Institute of Artificial Intelligence and Future Networks, Beijing Normal University(北京师范大学人工智能与未来网络研究院)
;
Faculty of Arts and Sciences, Beijing Normal University(北京师范大学文理学院)
;
Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学)