ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
VESTA: 一种全自动的LLM智能体场景生成与安全评估框架
Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
机构
*
BrainCog AI Lab, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所类脑人工智能实验室)
;
Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究院)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京市安全人工智能与超级对齐重点实验室)
;
School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)
;
Long-term AI(长期人工智能)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
"LLM Agent Performance" Is Not a Single Evaluation Target
基于LLM的智能体评估统一框架的必要性
Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
MEDLEY-BENCH:在AI元认知中规模购买评估但不控制
Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
机构
*
Department of Clinical Science, Intervention and Technology (CLINTEC), Karolinska Institutet(临床科学、干预与技术部门(CLINTEC),Karolinska研究所)
;
Department of Clinical Physiology, Karolinska University Hospital(临床生理学部门,Karolinska大学医院)
;
Department of Biomedical Engineering and Health Systems, KTH Royal Institute of Technology(生物医学工程与健康系统部门,KTH皇家理工学院)
;
Department of Textile Technology, University of Borås(纺织技术部门,Borås大学)
;
Department of Medical Technologies, Karolinska University Hospital(医学技术部门,Karolinska大学医院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Westlake University(西湖大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.LG
CommentsAccepted by Transactions on Machine Learning Research. (32 pages, 12 figures.) This version refines the paper structure, adds experimental results. Project page: https://spherelab.ai/SGP-Gen/
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
GraphVerse:面向多模态大语言模型的综合性视觉图推理基准
Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan
机构
*
New York University(纽约大学)
;
Sensetime Research(商汤科技研究院)
;
Tsinghua University(清华大学)
;
New York University Shanghai(上海纽约大学)
;
University of Georgia(佐治亚大学)
;
The Hong Kong Polytechnic University(香港理工大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract)
Georeferencing Non-Gazetteered Place Names using Biological Specimen Records
利用生物标本记录对非 gazetteer 地名进行地理配准
Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones
机构
*
School of Mathematical and Computational Sciences, Massey University(梅西大学数学与计算科学学院)
;
Joint Centre for Disaster Research, Massey University(梅西大学灾害研究联合中心)
;
School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院)
Comments5 pages, 1 figure, 1 table. Accepted at the FoRMA workshop (Foundation Models in the RO-MAN Age: Responsible Development for Social Robotics) at IEEE RO-MAN 2026, Kitakyushu, Japan. Workshop homepage: https://sites.google.com/cam.ac.uk/forma/