AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
AgenticEval: 向大型语言模型的代理和自演化安全评估迈进
Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li, Xibang Yang, Yan Teng, Xingjun Ma, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
;
East China Normal University(华东师范大学)
机构
*
Research Intern, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系研究实习生)
;
Ph.D. Student, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系博士生)
;
Undergraduate Student, Aerospace Program, University of California, Berkeley(加州大学伯克利分校航空航天项目本科生)
;
Full Professor, Department of Electrical Engineering and Computer Science, University of California, Berkeley(加州大学伯克利分校电气工程与计算机科学系教授)
;
Ph.D. Student, Department of Computer Science, George Washington University(乔治华盛顿大学计算机科学系博士生)
;
Associate Professor, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系副教授)
Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control
面向安全的语言理解系统在空交通管制中的评估
Yujing Chang, Yash Guleria, Duc-Thinh Pham, Nhut-Huy Pham, Ningli Wang, Vu N. Duong, Sameer Alam
机构
*
ATMRI, Nanyang Technological University (NTU), Singapore(航空交通管理研究所,南洋理工大学(NTU),新加坡)
;
School of Management, Indian Institute of Technology Mandi, India(管理学院,印度理工学院曼迪分校,印度)
;
Centre of AI Research, VinUniversity, Vietnam(人工智能研究中心,文大学,越南)
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance
迈向稳定的值对齐:引入独立模块以实现一致的价值引导
Wenhao Chen, Sirui Sun, Shengyuan Bai, Guojie Song
机构
*
School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室)
Document Retrieval Augmented Fine-Tuning (DRAFT) for safety-critical software assessments
用于安全关键软件评估的文档检索增强微调(DRAFT)
Regan Bolton, Mohammadreza Sheikhfathollahi, Simon Parkinson, Vanessa Vulovic, Gary Bamford, Dan Basher, Howard Parkinson
机构
*
Digital Transit Limited, 3M Buckley Innovation Centre, UK, HD1 3BD(数字交通有限公司,3M Buckley创新中心,英国,HD1 3BD)
;
Department of Computer Science, University of Huddersfield, UK, HD1 3DH(计算机科学系,赫德瑟菲尔德大学,英国,HD1 3DH)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Claw-Eval: 向可信的自主代理评估迈进
Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
The University of Hong Kong(香港大学)
机构
*
LAAS-CNRS(法国图卢兹Laas--cnrs研究所)
;
University of Toulouse(图卢兹大学)
;
Thales(泰勒斯)
;
ONERA(法国国家航空器研究与测试中心)
;
Espace-Dev, IRD, Université de Montpellier(Espace-Dev, 法国国家农业与食品研发机构, 图卢兹大学)
;
Universidade Federal do Rio Grande do Norte(巴西北里奥格兰德联邦大学)