ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
ClinHallu: 用于诊断医学多模态大语言模型推理中阶段式幻觉的基准
Sicheng Yang, Hangjie Yuan, Wenjun Zhang, Jinwang Wang, Yichen Qian, Weihua Chen, Fan Wang, Lei Zhu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
DAMO Academy, Alibaba Group(阿里巴巴达摩院)
;
Hupan Lab(湖畔实验室)
;
Zhejiang University(浙江大学)
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
LLM能否准确评分医学诊断和临床推理?
Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
机构
*
Wits MIND Institute, University of the Witwatersrand, Johannesburg, South Africa(维特士心理研究所,沃斯兰德大学,约翰内斯堡,南非)
;
Grai Labs, Cape Town, South Africa(格雷实验室,开普敦,南非)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(南非医学研究理事会疫苗和传染病分析研究组,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Internal Medicine, Charlotte Maxeke Johannesburg Academic Hospital, and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(内科学系,查理·马克斯凯约翰内斯堡学术医院,以及健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Department of Paediatrics and Child Health, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa(儿科学与儿童健康系,健康科学学院,沃斯兰德大学,约翰内斯堡,南非)
;
Wits MIND Institute, University of the Witwatersrand, Johannesbu(维特士心理研究所,沃斯兰德大学,约翰内斯堡)
Comments9 pages main text, 31 pages total (including references and appendix). 5 figures, 16 tables. Preprint under review. Code and data will be made available upon publication
机构
*
Department of Computer Science and Information Engineering, National Taiwan University(国立台湾大学计算机科学与资讯工程系)
;
National Taiwan University AI Center of Research Excellence(国立台湾大学人工智能研究中心)
Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs
条件 Vendi 分数:生成式 AI 模型和 LLM 的提示感知多样性评估
Mohammad Jalali, Azim Ospanov, Amin Gohari, Farzan Farnia
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)
;
Department of Information Engineering, The Chinese University of Hong Kong(信息工程系,香港中文大学)
专题命中
安全评测
:alignment(abstract);分类 cs.AI、cs.LG
AI总结
针对文本提示引导的生成模型,提出条件 Vendi 和条件 RKE 分数,通过条件熵分离模型自身多样性,并证明收敛性及在多个任务中恢复真实多样性排序。
Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation
激活引导引发突现失调:一项更全面的评估
Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu
机构
*
Nanyang Technological University(南洋理工大学)
;
Sun Yat-sen University(中山大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals
机构
*
School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院)
;
Sirindhorn International Institute of Technology, Thammasat University(泰国朱拉隆梭国际技术学院)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Deep Interdisciplinary Intelligence Lab (DI2 Lab)(深度跨学科智能实验室(DI2 Lab))
Video Understanding by Design: How Datasets Shape Video Models
通过设计理解视频:数据集如何塑造视频模型
Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao
机构
*
School of Engineering and Built Environment, Electrical and Electronic Engineering, Griffith University(工程与建筑环境学院,电气与电子工程学院,格里菲斯大学)
;
School of Computer Science and Engineering, University of New South Wales(计算机科学与工程学院,新南威尔士大学)