机构
*
Department of Radiology, Keio University Hospital, Tokyo, Japan(Keio大学医院放射科)
;
Division of Infectious Diseases and Infection Control, Keio University Hospital, Tokyo, Japan(Keio大学医院感染病学与感染控制科)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
机构
*
Guangdong Key Laboratory of Intelligent Transportation System, School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University(广东智能交通系统重点实验室,智能系统工程学院,中山大学深圳校区)
;
University of Amsterdam(阿姆斯特丹大学)
;
Aarhus University(阿arhus大学)
;
Beijing Institute of Artificial Intelligence, Beijing University of Technology(北京人工智能研究院,北京工业大学)
;
School of Information and Engineering, Chang’an University(信息工程学院,长安大学)
;
Key Laboratory of Road and Traffic Engineering of the Ministry of Education, Tongji University(交通工程教育部长实验室,同济大学)
;
Key Laboratory of Intelligent Transportation Technology and System, School of Transportation Science and Engineering, Beihang University(智能交通技术与系统重点实验室,交通运输科学与工程学院,北航)
;
Jiangsu Key Laboratory of Urban ITS, Jiangsu Province Collaborative Innovation Center of Modern Urban Traffic Technologies, School of Transportation, Southeast University(江苏城市ITS重点实验室,江苏省现代城市交通技术协同创新中心,交通学院,东南大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG
CommentsAccepted by Expert Systems with Applications
Is 'Hope' a person or an idea? A pilot benchmark for NER: comparing traditional NLP tools and large language models on ambiguous entities
Payam Latifi
机构
*
Department of Humanities University of Turin(人文学院 特林姆大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
Comments14 pages, 9 figures, 2 tables. This is a pilot study evaluating six NER systems -- three traditional tools (NLTK, spaCy, Stanza) and three LLMs (Gemini-1.5-flash, DeepSeek-V3, Qwen-3-4B) -- on a small, ambiguity-rich dataset of 119 tokens. The annotated dataset, prompts are provided in appendices for full reproducibility. All experiments were conducted on 14 May 2025
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
Mike A. Merrill, Akshay Paruchuri, Naghmeh Rezaei, Geza Kovacs, Javier Perez, Yun Liu, Erik Schenck, Nova Hammerquist, Jake Sunshine, Shyam Tailor, Kumar Ayush, Hao-Wei Su, Qian He, Cory Y. McLean, Mark Malhotra, Shwetak Patel, Jiening Zhan, Tim Althoff, Daniel McDuff, Xin Liu
机构
*
Google Research(谷歌研究)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
Comments53 pages, 7 main figures, 2 main tables, accepted to Nature Communications
机构
*
Dept. of Computer Engineering, Pune Institute of Computer Technology, Pune, India(计算机工程系,普那计算机技术研究所,普那,印度)
;
Indian Institute of Technology Madras, Chennai, India(印度理工学院马德拉斯学院,钦奈,印度)
;
L3Cube Labs, Pune, Maharashtra, India(L3Cube实验室,普那,马哈拉施特拉邦,印度)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);pretraining(abstract);分类 cs.CL、cs.LG
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
Wonjun Jeong, Dongseok Kim, Taegkeun Whangbo
机构
*
Department of Computer Engineering(计算机工程系)
;
Gachon University(高城大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
CommentsComments: 34 pages, 1 figure. v2: All "Consequence." statements in the Theoretical Analysis section relabeled as "Corollary."; duplicated values in Table 20 (previously identical to Table 15) corrected