机构
*
Centre for Industrial Software, University of Southern Denmark(丹麦南部大学工业软件中心)
;
Department of Electronics, Information and Bioengineering, Politecnico di Milano(米兰理工学院电子、信息与生物工程系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback
基于大语言模型判官的多维行为评估:用于代理股票预测系统的闭环强化学习反馈
Mohammad Al Ridhawi, Mahtab Haj Ali, Hussein Al Osman
机构
*
School of Electrical Engineering and Computer Science(电气工程与计算机科学学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG
Exploring Lightweight Large Language Models for Court View Generation
探索用于法院视图生成的轻量级大语言模型
Zhitian Hou, Tianyong Hao, Nanli Zeng, Zhixiong Chao, Kun Zeng
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
School of Computer Science, South China Normal University(华南师范大学计算机学院)
;
China Mobile Internet Co., Ltd.(中国移动互联网有限公司)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.CL、cs.AI
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Birmingham(伯明翰大学)
;
Zhejiang University(浙江大学)
;
East China Normal University(华东师范大学)
;
Alibaba Group(阿里巴巴集团)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
在百万笔记规模上系统评估LLM重新表述的合成临床笔记质量
Jinghui Liu, Sarvesh Soni, Anthony Nguyen
机构
*
Australian e-Health Research Centre, CSIRO, Australia(澳大利亚电子健康研究中心,CSIRO,澳大利亚)
;
National Library of Medicine, National Institutes of Health, USA(国家医学图书馆,国立卫生研究院,美国)
专题命中
评测与基准
:LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
对多模态大语言模型评分者的审计:临床顺序评分中的中间倾向偏差
Jiaqing Zhang, Sandeep Elluri, Bhanu Cherukuvada, Yonah Joffe, Jessica Sena, Miguel Contreras, Scott Siegel, Subhash Nerella, Catherine Price, Parisa Rashidi
机构
*
Department of Electrical & Computer Engineering(电气与计算机工程系)
;
Department of Computer and Information Science and Engineering(计算机与信息科学与工程系)
;
Department of Clinical and Health Psychology(临床与健康心理学系)
;
Department of Biomedical Engineering(生物医学工程系)
专题命中
评测与基准
:LLM(title,summary_cn);large language model(abstract);language model(abstract)
机构
*
Northeastern University(东北大学)
;
University of Southern California(南加州大学)
;
Stony Brook University(石溪大学)
;
Independent Researcher(独立研究者)
;
Ohio State University(俄亥俄州立大学)
;
University of Notre Dame(Notre Dame 大学)
;
Columbia University(哥伦比亚大学)
专题命中
评测与基准
:LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL
HalluScore: Large Language Model Hallucination Question Answering Benchmark
HalluScore: 大语言模型幻觉问答基准
Aisha Alansari, Hamzah Luqman
机构
*
Department of Information and Computer Science, King Fahd University of Petroleum and Minerals(国王法赫德石油与矿物大学信息与计算机科学系)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.CL
机构
*
School of Economics and Management, East China Normal University(东华大学经济管理学院)
;
School of Information Management, Wuhan University(武汉大学信息管理学院)
;
China Academic Degrees & Graduate Education Development Center(中国学位与研究生教育发展中心)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL