BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
评估两用生物学环境中的校准拒绝和安全有用性
Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
Comments23 pages, 14 figures. Major revision merging the follow-up "Standing Invariants at Scale" into this paper per arXiv moderation: the machine-checked action-safety core preserved across six further releases, six new invariant families with teeth, and real-hardware self-improvement; suite grown from 122 to 563 tests. Software, gate suite, run artifacts, and TLA+ spec are open source
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
需要多少次迭代才能突破限制?多轮LLM评估中的动态预算分配
Shai Feldman, Yaniv Romano
机构
*
Department of Computer Science(计算机科学系)
;
Technion, Israel(技术ion, 以色列)
;
Departments of Electrical and Computer Engineering and of Computer Science(电气与计算机工程系和计算机科学系)
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
大型语言模型的红队框架:以忠实性评估为例
Abrar Alotaibi, Raed Mughus, Moataz Ahmed
机构
*
King Fahd University of Petroleum & Minerals(法赫德国王石油矿产大学)
;
Imam Abdulrahman Bin Faisal University(伊玛目阿卜杜勒拉赫曼·本·费萨尔大学)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence(沙特数据与人工智能局-法赫德国王石油矿产大学人工智能联合研究中心)
Comments26 pages, 5 figures, 9 tables. v2: cite the released Mental Spaces Corpus (dataset DOI); switch to ACL bibliography style; minor copy-editing. No change to results
Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It
重新定义AI失控:它是什么,如何拥有,如何失去
Ze Shen Chin, Maurice Chiodo, Dennis Müller, Coleman Snell
机构
*
Oxford Martin AI Governance Initiative AI Standards Lab(牛津马丁人工智能治理倡议人工智能标准实验室)
;
Centre for the Study of Existential Risk, University of Cambridge(存在风险研究中心,剑桥大学)
;
Institute of Mathematics Education, University of Cologne(数学教育研究所,科隆大学)
;
Cornell University(康奈尔大学)
A Self-Supervised Framework for Space Object Behaviour Characterisation
一种用于空间物体行为特征刻画的自监督框架
Ian Groves, Andrew Campbell, James Fernandes, Diego Ramírez Rodríguez, Paul Murray, Massimiliano Vasile, Victoria Nockles
机构
*
Defence AI Research Centre, Defence & National Security, The Alan Turing Institute(国防人工智能研究中心,国防与国家安全,艾伦·图灵研究所)
;
Department of Electronic and Electrical Engineering, University of Strathclyde(电子与电气工程系,斯特拉斯克莱德大学)
;
GMV
;
Aerospace Centre of Excellence, University of Strathclyde(航空航天卓越中心,斯特拉斯克莱德大学)