Comments17 pages. Published in the Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2) at ACL 2025
Journal refProceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2), pages 320-336, Association for Computational Linguistics, 2025
CommentsICAHS, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Comments19 pages, 5 figures, 3 tables. Benchmark paper introducing SPLIT for evaluating empathy, linguistic naturalness, and cultural grounding in English and Ukrainian LLM responses
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
CORDA:面向大语言模型的以伤害为核心的分层道德推理基准
Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
机构
*
University of the Witwatersrand(威特沃特斯兰德大学)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
一种智能体AI框架克服了大型语言模型在眼底照相青光眼检测中的基本局限
Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
Large language models improve physician accuracy but lead to false reliance
大型语言模型可提高医师准确率,但会导致错误依赖
Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models
OntoLearner: 一个用于大语言模型本体学习的模块化Python库
Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei, Nandana Mihindukulasooriya, Sören Auer
机构
*
TIB – Leibniz Information Centre for Science and Technology(TIB – 莱布尼茨科学与技术信息中心)
;
L3S Research Center, Leibniz University of Hannover(L3S研究中心,莱布尼茨汉诺威大学)
;
IBM Research(IBM研究院)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
机构
*
Division of Radiology and Biomedical Engineering, Graduate School of Medicine, The University of Tokyo(放射医学与生物医学工程系,东京大学医学研究生院)
;
Department of Computational Diagnostic Radiology and Preventive Medicine, The University of Tokyo Hospital(计算诊断放射学与预防医学系,东京大学医院)
;
Department of Radiology, Kanazawa University Hospital(金泽大学医院放射科)
;
Faculty of Medicine, The University of Tokyo(东京大学医学系)
;
Department of Radiology, School of Medicine, Jichi Medical University(立命馆大学医学系放射科)
;
Department of Radiology, The University of Tokyo Hospital(东京大学医院放射科)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI