Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
意外漏洞:影响微调的因子
Punya Syon Pandey, Samuel Simko, Kellin Pelrine, Zhijing Jin
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
ETH Zürich(苏黎世联邦理工学院)
;
FAR AI(FAR人工智能公司)
;
McGill University(麦吉尔大学)
;
MILA(蒙特利尔人工智能研究院)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(图宾根德国智能系统研究所)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
基于能力的LLM红队测试规模趋势
Alexander Panfilov, Paul Kassianik, Maksym Andriushchenko, Jonas Geiping
机构
*
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Tübingen AI Center(图宾根人工智能中心)
;
Foundation AI – Cisco Systems Inc.(AI基础研究机构——思科系统公司)
;
EPFL(苏黎世联邦理工学院)
Trust in LLM-controlled Robotics: a Survey of Security Threats, Defenses and Challenges
对LLM控制的机器人信任:安全威胁、防御和挑战的综述
Xinyu Huang, Shyam Karthick V B, Taozhao Chen, Mitch Bryson, Thomas Chaffey, Huaming Chen, Kim-Kwang Raymond Choo, Ian R. Manchester
机构
*
School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院)
;
Australian Centre for Robotics and School of Aerospace, Mechanical and Mechatronic Engineering, The University of Sydney(悉尼大学机器人中心及航空航天、机械与机电工程学院)
;
Department of Information Systems and Cybersecurity, University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校信息系统与网络安全系)
;
School of Engineering and Natural Sciences, University of Iceland(冰岛大学工程与自然科学学院)