机构
*
Central South University(中南大学)
;
Tsinghua University(清华大学)
;
Nanjing University(南京大学)
;
Suzhou Aerospace Information Research Institute(苏州空天信息研究院)
;
McGill University(麦吉尔大学)
机构
*
Institute of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
Information Research Center of Military Science, PLA Academy of Military Science(军事科学院军事科学信息研究中心)
机构
*
Nanjing University(南京大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models
推理下的校准漂移:思维链预算如何导致大型语言模型过度自信
Prakul Sunil Hiremath, Harshit R. Hiremath
机构
*
Department of Computer Science and Engineering, Visvesvaraya Technological University, Belagavi(维斯瓦拉亚科技大学计算机科学与工程系,贝拉加维)
;
Department of Computer Science and Business System, SG Balekundri Institute of Technology, Belagavi(SG巴莱昆德里理工学院计算机科学与商业系统系,贝拉加维)
Comments31 pages, 4 figures, 3 tables. Introduces Calibration Drift Under Reasoning (CDUR) with theoretical analysis and preliminary experiments; includes CABStop; code and data available
MammoExpert: Benchmarking Chain-of-Thought Reasoning in Mammography Diagnosis
MammoExpert:乳腺X线诊断中的思维链推理基准测试
Di Dai, Bo Liu, Youcheng Li, Haojun Yu, Zhouhang Bian, Quanlin Wu, Dong Wang, Sichen Meng, Hongye Xuan, Zijie Lan, Shenda Hong, Liwei Wang
机构
*
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室)
;
School of Computer Science and Engineer, Beijing University of Aeronautics and Astronautics(北京航空航天大学计算机科学与工程学院)
;
Center for Data Science, Peking University(北京大学数据科学中心)
;
Yizhun co. ltd(医准有限公司)
;
International School, Beijing University of Post and Telecommunications(北京邮电大学国际学院)
;
School of Public Health, University of Michigan, Ann Arbor(密歇根大学安娜堡分校公共卫生学院)
;
Future Technology College, Xi'an Jiaotong University(西安交通大学未来技术学院)
;
Peking University(北京大学)
Han Huang, Hao Wang, Mengqi Zhang, Shu Wu, Qiang Liu, Liang Wang
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院自动化研究所模式识别国家重点实验室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Shandong University(山东大学)
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias