ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
ClinHallu: 用于诊断医学多模态大语言模型推理中阶段式幻觉的基准
Sicheng Yang, Hangjie Yuan, Wenjun Zhang, Jinwang Wang, Yichen Qian, Weihua Chen, Fan Wang, Lei Zhu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
DAMO Academy, Alibaba Group(阿里巴巴达摩院)
;
Hupan Lab(湖畔实验室)
;
Zhejiang University(浙江大学)
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
LLM能否准确评分医学诊断和临床推理?
Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
机构
*
Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非)
;
Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)
Digital Twin Driven Textile Classification and Foreign Object Recognition in Automated Sorting Systems
数字孪生驱动的自动化分拣系统中的纺织品分类与异物识别
Serkan Ergun, Tobias Mitterer, Hubert Zangl
机构
*
Institute of Smart Systems Technologies(智能系统技术研究所)
;
University of Klagenfurt(克雷格弗特大学)
;
AAU SAL USE Laboratory(AAU SAL USE实验室)
;
Silicon Austria Labs(硅 Austria 实验室)
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
DLawBench: 通过多轮法律咨询评估大语言模型
Li Zhang, Yuzhen Shi, Yiran Hu, Jingwen Zhang, Wenbo Lv, Yubo Ma, Wei Wang, Rongyao Shi, Yuanyang Qiu, Xinran Xu, Yuemeng Qi, Linlin Miao, Jaromir Savelka, Yun Liu, Kevin Ashley, Bing Zhao, Hu Wei, Lin Qu
机构
*
University of Pittsburgh(匹兹堡大学)
;
Alibaba Group(阿里巴巴集团)
;
University of Waterloo(滑铁卢大学)
;
Shandong University(山东大学)
;
Skylenage
;
China University of Political Science and Law(中国政法大学)
;
Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)
;
Fordham University(福特汉姆大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Tsinghua University(清华大学)
机构
*
Missouri University of Science and Technology(密苏里科技大学)
;
University of South Florida(佛罗里达州立大学)
;
Visa Inc.(Visa公司)
;
George Mason University(乔治·马歇尔大学)
CommentsThis is the author's accepted version of the paper accepted to appear at IEEE AIIoT 2025. The final version will be available via IEEE Xplore. \c{opyright}2025 IEEE. Personal use of this material is permitted