机构
*
School of Physics and Astronomy, The University of Edinburgh(物理学与天文学学院,爱丁堡大学)
;
Glasgow College, University of Electronic Science and Technology of China(电子科技大学成都学院)
;
School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(机械与电子工程学院,电子科技大学)
;
Department of Systems Engineering, City University of Hong Kong(系统工程系,香港城市大学)
S2SServiceBench: A Multimodal Benchmark for Last-Mile S2S Climate Services
S2SServiceBench:一个多模态基准用于最后一公里S2S气候服务
Chenyue Li, Wen Deng, Zhuotao Sun, Mengxi Jin, Hanzhe Cui, Han Li, Shentong Li, Man Kit Yu, Ming Long Lai, Yuhao Yang, Mengqian Lu, Binhang Yuan
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Nanjing University of Information Science and Technology(南京信息工程大学)
;
Beijing Normal University(北京师范大学)
机构
*
National University of Science and Technology POLITEHNICA Bucharest, Faculty of Automatic Control and Computers(波兰技术大学布加勒斯特分校)
;
Technical University of Munich(慕尼黑技术大学)
;
National Institute for Research & Development in Informatics - ICI Bucharest(信息研究所-布加勒斯特)
;
Paris 1 Panthéon-Sorbonne University(巴黎1大学)
;
University of Bucharest(布加勒斯特大学)
机构
*
Institute of Trustworthy Embodied AI(可信具身人工智能研究所)
;
Institute of Technology Ethics for Human Future(人类未来技术伦理研究所)
;
School of Philosophy(哲学学院)
;
School of Life Sciences(生命科学学院)
;
Ethics Committee of Zhongshan Hospital(中山医院伦理委员会)
;
Department of Cardiology, Zhongshan Hospital of Fudan University, Institute of Cardiovascular Diseases, National Clinical Research Centre for Interventional Medicine(复旦大学中山医院心内科、心血管疾病研究所、介入医学临床研究中心)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
CommentsThe authors are withdrawing this preprint as it was submitted prematurely without the final approval of all collaborating institutions. We apologize for any inconvenience
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
医疗与医学中的代理AI:一个七维分类法用于评估基于大语言模型的代理
Shubham Vatsal, Harsh Dubey, Aditi Singh
机构
*
Department of Computer Science, New York University, CIMS, New York, USA(纽约大学计算机科学系,CIMS,纽约,美国)
;
Department of Computer Science, Cleveland State University, Cleveland, USA(克利夫兰州立大学计算机科学系,克利夫兰,美国)
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
Tianjin University(天津大学)
;
Peking University(北京大学)
;
Zhejiang University(浙江大学)
;
Beijing Institute of General Artificial Intelligence(北京一般人工智能研究院)
OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models
OMGEval: 一种开放的多语言生成评估基准用于大型语言模型
Yang Liu, Meng Xu, Shuo Wang, Liner Yang, Haoyu Wang, Zhenghao Liu, Cunliang Kong, Yun Chen, Yang Liu, Maosong Sun, Erhong Yang
机构
*
Beijing Language and Culture University(北京语言大学)
;
Tsinghua University(清华大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Northeastern University(东北大学)
;
Shanghai University of Finance and Economics(上海金融学院)
MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
MasalBench: 一个用于评估大语言模型对波斯谚语的上下文和跨文化理解的基准
Ghazal Kalhor, Behnam Bahrak
机构
*
School of Electrical and Computer Engineering, College of Engineering, University of Tehran(德黑兰大学电气与计算机工程学院,工程学院)
;
Tehran Institute for Advanced Studies, Khatam University(德黑兰高级研究 institute,Khatam 大学)
机构
*
Department of Computer and Network Engineering, College of Information Technology, United Arab Emirates University(计算机与网络工程系,信息科技学院,阿拉伯联合酋长国大学)
;
Khalifa University of Science and Technology(科学与技术喀拉奇大学)
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
MMR-Bench: 多模态大语言模型路由的综合基准
Haoxuan Ma, Guannan Lai, Han-Jia Ye
机构
*
School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学)
;
National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学)