A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
基于决策理论的隐写术形式化及其在大语言模型监控中的应用
Usman Anwar, Julianna Piskorz, David D. Baek, David Africa, Jim Weatherall, Max Tegmark, Christian Schroeder de Witt, Mihaela van der Schaar, David Krueger
机构
*
University of Cambridge(剑桥大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
UK AI Safety Institute(英国人工智能安全研究所)
;
University of Oxford(牛津大学)
;
Mila, University of Montreal(蒙特利尔大学米尔人工智能实验室)
Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, Robert Trager
机构
*
Oxford Martin AI Governance Initiative, University of Oxford(牛津大学人工智能治理倡议)
;
MIT Computer Science and Artificial Intelligence Laboratory, MIT(麻省理工学院计算机科学与人工智能实验室)
;
MIT Future Tech(麻省理工学院未来技术)
;
Stanford University(斯坦福大学)
;
Governance and Responsible AI Lab, Purdue University(普渡大学治理与负责任的人工智能实验室)
;
University of Toronto(多伦多大学)
;
Mercatus Center, George Mason University(乔治·马歇尔大学麦卡锡中心)
;
Vilnius University(维尔纽斯大学)
;
Vijil
;
SaferAI
;
AI Standards Lab(人工智能标准实验室)
;
The Future Society(未来社会)
;
Concordia AI(康科德人工智能)
;
Pivotal Research
;
Center for AI Risk Management & Alignment(人工智能风险管理和对齐中心)
;
UC Berkeley Center for Long-Term Cybersecurity(伯克利大学长期网络安全中心)
;
Independent(独立)
Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
自主知识图谱探索与自适应广度-深度检索
Joaquín Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori, David A. Clifton, Luciano Del Corro, Marinka Zitnik
机构
*
Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系)
;
Departamento de Computación, FCEyN, Universidad de Buenos Aires(布宜诺斯艾利斯大学计算机系)
;
Department of Engineering Science, University of Oxford(牛津大学工程科学系)
;
Oxford Suzhou Centre for Advanced Research, University of Oxford(牛津大学苏州市先进研究中心)
;
ELIAS Lab, Departamento de Ingeniería, Universidad de San Andrés(圣安德鲁大学工程系ELIAS实验室)
;
Kempner Institute for the Study of Natural and Artificial Intelligence, Allston, MA, USA(自然与人工智能研究所,马萨诸塞州阿利斯顿)
;
Broad Institute of MIT and Harvard, Cambridge, MA, USA(MIT和哈佛大学Broad研究所)
;
Harvard Data Science Initiative, Cambridge, MA, USA(哈佛大学数据科学倡议)
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
多智能体安全中的开放挑战:迈向交互式AI智能体的安全系统
Christian Schroeder de Witt, Klaudia Krawiecka, Igor Krawczuk, Ben Hagag, William L. Anderson, Peter Belcak, Ben Bucknall, Xiaohong Cai, Ayush Chopra, Doron Cohen, Ron F. Del Rosario, Andis Draguns, Annie Gray, Keren Katz, Vasilios Mavroudis, Jaron Mink, Sumeet Ramesh Motwani, Jonathan Petit, Leif-Sebastian Rembeck, Chandler Smith, John Sotiropoulos, Steven Young, Sarah Scheffler, Mary Llewellyn
机构
*
Oxford Witt Lab, University of Oxford(牛津Witt实验室,牛津大学)
;
Department of Engineering Science, University of Oxford(牛津大学工程科学系)
;
Association for Computing Machinery (ACM)(计算机协会(ACM))
;
Independent(独立)
;
MATS Research(MATS研究)
;
CyLab Security & Privacy Institute, Carnegie Mellon University(CyLab安全与隐私研究所,卡内基梅隆大学)
;
Qualcomm Inc.(高通公司)
;
Oxford Martin AI Governance Initiative(牛津马丁人工智能治理倡议)
;
Carnegie Mellon University(卡内基梅隆大学)
;
MIT Media Lab(麻省理工媒体实验室)
;
SAP SE(SAP德国分公司)
;
OWASP GenAI Security Project - Agentic Security Initiative(OWASP生成式AI安全项目-代理安全倡议)
;
Contramont Research(Contramont研究)
;
The Alan Turing Institute(艾伦·图灵研究所)
;
Department of Economics, New York University(纽约大学经济系)
;
Zenity(Zenity公司)
;
King’s College London(伦敦国王学院)
;
Arizona State University(亚利桑那州立大学)
;
Torr Vision Group, University of Oxford(托尔视觉组,牛津大学)
;
Deep Cyber Ltd(Deep Cyber有限公司)