CommentsThis paper has been accepted by AAAI 2026. We update it for adding new evaluation results for ArxivRollBench-2025a and ArxivRollBench-2026a, with the evaluation of timly models like DeepSeekV4Pro, GPT-5.5, Claude-Opus-4.7, and so on. Source code: https://github.com/liangzid/ArxivRoll/ Online Leaderboard Website: https://arxivroll.moreoverai.com/
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
AgenticEval: 向大型语言模型的代理和自演化安全评估迈进
Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li, Xibang Yang, Yan Teng, Xingjun Ma, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
;
East China Normal University(华东师范大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.AI
Yahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank, Angel Hsing-Chi Hwang, Ruishan Liu
机构
*
Department of Computer Science, University of Southern California(南加州大学计算机科学系)
;
Department of Electrical and Computer Engineering, University of Southern California(南加州大学电气与计算机工程系)
;
Suzanne Dworak-Peck School of Social Work, University of Southern California(南加州大学苏兹安·德沃拉克-佩克社会工作学院)
;
Department of Psychiatry and the Behavioral Sciences, University of Southern California(南加州大学精神病学与行为科学系)
;
Annenberg School for Communication, University of Southern California(南加州大学安纳伯格通信学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
机构
*
Research Intern, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系研究实习生)
;
Ph.D. Student, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系博士生)
;
Undergraduate Student, Aerospace Program, University of California, Berkeley(加州大学伯克利分校航空航天项目本科生)
;
Full Professor, Department of Electrical Engineering and Computer Science, University of California, Berkeley(加州大学伯克利分校电气工程与计算机科学系教授)
;
Ph.D. Student, Department of Computer Science, George Washington University(乔治华盛顿大学计算机科学系博士生)
;
Associate Professor, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系副教授)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models
对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估
William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Bertan Ucar, Vitor D. de Moura, José O. Gomes
机构
*
Department of Industrial Engineering, Tsinghua University(清华大学工业工程系)
;
School of Social Sciences, Tsinghua University(清华大学社会科学部)
;
Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
Comments26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias
aLLoyM: A large language model for alloy phase diagram prediction
aLLoyM:一种用于合金相图预测的大型语言模型
Yuna Oikawa, Guillaume Deffrennes, Taichi Abe, Ryo Tamura, Koji Tsuda
机构
*
Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究院)
;
University Grenoble Alpes, CNRS, Grenoble INP, SIMaP(格勒诺布尔阿尔卑斯大学、CNRS、格勒诺布尔INP、SIMaP)
;
Research Center for Structural Materials, National Institute for Materials Science(材料研究所结构材料研究中心)
;
Center for Basic Research on Materials, National Institute for Materials Science(材料研究所基础材料研究中心)
;
RIKEN Center for Advanced Intelligence Project(理化学研究所先进人工智能项目中心)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
机构
*
University of Macau(澳门大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong(香港城市大学)
;
Seoul National University(首尔国立大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Stevens Institute of Technology(史蒂文斯理工学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
A Systematic Approach for Large Language Models Debugging
大语言模型调试的系统方法
Basel Shbita, Anna Lisa Gentile, Bing Zhang, Sungeun An, Shailja Thakur, Shubhi Asthana, Yi Zhou, Saptha Surendran, Farhan Ahmed, Rohan Kulkarni, Yuya Jeremy Ong, Chad DeLuca, Hima Patel
机构
*
IBM Research(IBM研究院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI