PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
PortBench: 一种相关性感知的、全流水线的LLM驱动投资组合管理基准
Yuxuan Zhao, Sijia Chen, Ningxin Su
机构
*
Yantai Research Institute of Harbin Engineering University(哈尔滨工程大学烟台研究院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
Med-CRAFT:通过知识图谱遍历自动构建可解释的多跳视频工作负载
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
Beijing Institute of Technology(北京理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care
RESPClinBench:呼吸专科医疗中的多模态临床决策与纵向疾病管理基准测试
Mouxiao Bian, Zhi Chen, Ruiyao Chen, Lu Lu, Hengrui Liang, Chaoyi Huang, Yiluo Lin, Jingru Ding, Yun Zhong, Yueming Su, Jie Xu
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Macau University of Science and Technology(澳门科技大学)
;
First Affiliated Hospital of Guangzhou Medical University(广州医科大学附属第一医院)
;
Guangzhou Institute of Respiratory Health(广州呼吸健康研究院)
Clinician input steers AI toward accurate and harmful recommendations
临床输入引导前沿AI模型做出准确和有害的决策
Ivan Lopez, Selin S. Everett, Bryan J. Bunning, April S. Liang, Dong Han Yao, Shivam C. Vedak, Kameron C. Black, Sophie Ostmeier, Stephen P. Ma, Emily Alsentzer, Jonathan H. Chen, Akshay S. Chaudhari, Eric Horvitz
机构
*
Ant Group(蚂蚁集团)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Xi'an Polytechnic University(西安理工大学)
An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness
基于临床数据的AI模型更新风险实证评估:稳定性、任意性与公平性
Ioannis Bilionis, Ricardo C. Berrios, Luis Fernandez-Luque, Carlos Castillo
机构
*
Spanish Ministry of Science and Innovation(西班牙科学与创新部)
;
Department of Research and Universities of the Government of Catalonia(加泰罗尼亚政府研究与大学部门)
;
MCIN/AEI /10.13039/501100011033(MCIN/AEI)
;
Maria de Maeztu Units of Excellence Programme(玛丽亚·德·玛埃斯特乌卓越计划)
机构
*
China University of Petroleum-Beijing at Karamay(中国石油大学(北京)克拉玛依校区)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Tianjin University(天津大学)
CommentsThis paper has been withdrawn due to a complete pivot in research direction and methodology. We have shifted from the atmospheric domain application to a general methodological direction. A new version of this research will be submitted separately
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
临床医生级别的一致性缺乏临床谨慎:LLM评估者在医学AI基准测试中的局限性
William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian Löns, Ronald Böck, Sebastian Fudickar
机构
*
University of Luebeck(吕贝克大学)
;
University of Tübingen(图宾根大学)
;
University Hospital Schleswig-Holstein(石勒苏益格-荷尔斯泰因大学医院)
;
University Hospital Würzburg(维尔茨堡大学医院)
;
Charité – Universitätsmedizin Berlin(柏林夏里特医学院)
;
Genie Enterprise Deutschland GmbH
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断
Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng
机构
*
University of Liverpool(利物浦大学)
;
Institute of Life Course & Medical Sciences(生命课程与医学科学研究院)
;
Department of Eye and Vision Sciences(眼科与视觉科学系)
;
Computer Science Department(计算机科学系)
;
Cardiovascular & Metabolic Medicine(心血管与代谢医学)
;
Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)