Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Agent-ValueBench: 一个全面的评估代理价值观的基准
Haonan Dong, Qiguan Feng, Kehan Jiang, Haoran Ye, Xin Zhang, Guojie Song
机构
*
State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
;
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
Peking University School of Software and Microelectronics(北京大学软件与微电子学院)
;
Peking University School of Psychological and Cognitive Sciences(北京大学心理与认知科学学院)
;
Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学机器感知重点实验室)
机构
*
Northeastern University(东北大学)
;
University of Pécs(佩奇大学)
;
University of Michigan(密歇根大学)
;
AT&T Chief Data Office(AT&T首席数据办公室)
;
Indian Institute of Management Bangalore(班加罗尔印度管理学院)
机构
*
Interdepartmental Program in Computational Biology and Bioinformatics, Yale University(耶鲁大学计算生物学与生物信息学联合计划)
;
Department of Biostatistics, Yale University(耶鲁大学生物统计学系)
;
Broad Institute of MIT and Harvard(哈佛大学与麻省理工学院联合Broad研究所)
;
Center for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡医学学校生物医学数据科学中心)
;
Sport and Exercise Medicine Service, KK Women’s and Children’s Hospital Training Program, Duke-NUS Medical School(杜克-新加坡医学学校KK妇女儿童医院运动与医学服务培训项目)
;
Training Program, Duke-NUS Medical School(杜克-新加坡医学学校培训项目)
;
Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系)
;
Department of Complexity Science and Engineering, The University of Tokyo(东京大学复杂科学与工程系)
;
Center for Advanced Intelligence Project, RIKEN(日本理化学研究所高级智能项目中心)
;
NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所)
;
Department of Biostatistics and Bioinformatics, Duke University(杜克大学生物统计学与生物信息学系)
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
Department of Genetics, Yale University(耶鲁大学遗传学系)
;
Wu Tsai Institute, Yale University(耶鲁大学吴天教授研究所)
;
Department of Biomedical Informatics and Data Science, Yale University(耶鲁大学生物医学信息学与数据科学系)
MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study
MATRA:代理AI系统的攻击面建模——OpenClaw案例研究
Tim Van hamme, Thomas Vissers, Javier Carnerero-Cano, Mario Fritz, Emil C. Lupu, Lieven Desmet, Dinil Mon Divakaran
机构
*
DistriNet, KU Leuven(DistriNet,根特大学)
;
IBM Research(IBM研究院)
;
CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心)
;
Imperial College London(伦敦帝国理工学院)
;
A*STAR Institute for Infocomm Research(A*STAR信息与通信研究机构)
CommentsAccepted for presentation at the 5th International Workshop on Designing and Measuring Security in Systems with AI (DeMeSSAI 2026), co-located with the 11th IEEE European Symposium on Security and Privacy (EuroS&P 2026), Lisbon, Portugal, July 10, 2026
Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei, Junwen Miao, Huaisong Zhang, Songcheng Cai, Yubo Wang, Dongfu Jiang, Yuyu Zhang, Ping Nie, Wenhu Chen, Changqian Yu, Kelsey R. Allen
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(向量研究所)
;
Kolors Team, Kuaishou Technology(快手团队)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Waterloo(滑铁卢大学)
;
Etude AI
;
Tsinghua University(清华大学)
;
Georgia Institute of Technology(佐治亚理工学院)
MDGYM: Benchmarking AI Agents on Molecular Simulations
MDGYM:在分子模拟上评估AI代理的基准测试
Vinay Kumar, Satyendra Rajput, Mausam, N. M. Anoop Krishnan
机构
*
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院)
;
Department of Computer Science and Engineering, Indian Institute of Technology Delhi(印度理工学院德里计算机科学与工程系)
;
Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi(印度理工学院德里土木与环境工程系)
Commentsdue to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract appearing here is slightly shorter than that in the PDF file
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
AnomalyClaw:通过工具引导的反驳实现通用视觉异常检测代理
Xi Jiang, Yinjie Zhao, Zesheng Yang, Feng Zheng
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology (SUSTech), Shenzhen, China(南方科技大学计算机科学与工程系,深圳,中国)
;
School of EEE, Nanyang Technological University (NTU), Singapore(南洋理工大学电子工程学院,新加坡)
;
CFAR, Agency for Science, Technology and Research (A*STAR), Singapore(科技研究局(A*STAR)的CFAR,新加坡)
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Tsinghua University(清华大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Zhejiang University(浙江大学)
;
Nanyang Technological University(南洋理工大学)
Conformity Generates Collective Misalignment in AI Agents Societies
一致性在AI代理社会中产生集体偏离
Giordano De Marzo, Alessandro Bellina, Claudio Castellano, Viola Priesemann, David Garcia
机构
*
University of Konstanz(康斯坦茨大学)
;
Sony Computer Science Laboratories - Rome(索尼计算机科学实验室-罗马)
;
Sapienza University of Rome(罗马大学)
;
Max Planck Institute for Dynamics and Self-Organization(马克斯·普朗克动态与自组织研究所)
;
Institute for the Dynamics of Complex Systems, University of Gottingen(复杂系统动力学研究所,哥廷根大学)
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
一个评估自主AI代理结果驱动约束违反的基准
Miles Q. Li, Benjamin C. M. Fung, Martin Weiss, Pulei Xiong, Khalil Al-Hussaeni, Claude Fachkha
机构
*
McGill University(麦吉尔大学)
;
Tiptree Advanced Systems Corporation(蒂普里高级系统公司)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
National Research Council Canada(加拿大国家研究委员会)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
University of Dubai(迪拜大学)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
DeepTumorVQA: 一种分层的3D CT基准,用于分阶段评估医学视觉语言模型和工具增强代理
Yixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi, Boyan Wang, Liang He, Xinze Zhou, Sezgin Er, Ibrahim Ethem Hamamci, Zongwei Zhou, Alan Yuille
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
University of Bologna(博洛尼亚大学)
;
Istanbul Medipol University(伊斯坦布尔梅迪波尔大学)
;
Center for Biomolecular Nanotechnologies, Istituto Italiano di Tecnologia(生物分子纳米技术中心,意大利技术研究院)
;
The First Affiliated Hospital, Sun Yat-Sen University(中山大学第一附属医院)
;
Tongji University(同济大学)
Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving
超越自我博弈与规模:面向自动驾驶泛化的行为基准
Aron Distelzweig, Faris Janjoš, Andreas Look, Anna Rothenhäusler, Daniel Jost, Oliver Scheel, Raghu Rajan, Daphne Cornelisse, Eugene Vinitsky, Joschka Boedecker
机构
*
University of Freiburg(弗赖堡大学)
;
Bosch Center for Artificial Intelligence(博世人工智能中心)
;
Coburg University of Applied Sciences(科堡应用科学大学)
;
New York University(纽约大学)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
AgentCollabBench: 评估好代理为何会成为差合作者
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia, Tanzila Khan, Kainat Raisa Hossain, Nehaa Shri, Shubhrangshu Debsarkar, Humayra Tasnim, Gour Gupal Talukder Shawon, Debjoty Mitra, Sumaiya Ahmed Rani, Al Jami Islam Anik, Al Nafeu Khan
机构
*
University of Utah(犹他大学)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
University of Dhaka(达卡大学)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
University of Virginia(弗吉尼亚大学)
;
Rajshahi University of Engineering and Technology(拉贾加赫尔工程与技术大学)
;
Shahjalal University of Science and Technology(沙赫jalal科学与技术大学)
;
BRAC University(BRAC大学)
;
Islamic University of Technology(伊斯兰技术大学)
;
Comilla University(科摩拉大学)