Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Claudini: 自动研究发现针对LLM的最先进对抗攻击算法
Alexander Panfilov, Peter Romov, Igor Shilov, Yves-Alexandre de Montjoye, Jonas Geiping, Maksym Andriushchenko
机构
*
MATS ELLIS Institute(MATS ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Tübingen AI Center(图宾根人工智能中心)
;
Imperial College London(伦敦帝国理工学院)
机构
*
Department of Computer Science, University of Texas at El Paso(德克萨斯大学埃尔帕索分校计算机科学系)
;
School of Computing, Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校计算机学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
JC STEM Lab of Machine Learning and Computer Vision(机器学习与计算机视觉的JC STEM实验室)
;
Hong Kong Jockey Club Charities Trust(香港赛马会慈善信托基金)
;
Global STEM Professorship Scheme of the Hong Kong Special Administrative Region (SAR)(香港特别行政区(SAR)的全球STEM教授计划)
;
National Natural Science Foundation of China(中国国家自然科学基金委员会)
;
Northwestern Polytechnical University(西北工业大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
School of Electronics and Information(电子与信息学院)
;
Department of Electrical and Electronic Engineering(电气与电子工程系)
;
Hong Kong Institute of AI for Science(香港人工智能科学研究院)
;
City University of Hong Kong(香港城市大学)
机构
*
Department of Computer Science, University of Texas at El Paso(德克萨斯理工大学计算机科学系)
;
School of Computing, Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校计算机学院)
;
Hanyang University(翰阳大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
针对LLM编码代理技能生态系统的供应链污染攻击
Yubin Qu, Yi Liu, Tongcheng Geng, Gelei Deng, Yuekang Li, Leo Yu Zhang, Ying Zhang, Lei Ma
机构
*
Griffith University(格里菲斯大学)
;
The State Information Center(国家信息中心)
;
Nanyang Technological University(南洋理工大学)
;
University of New South Wales(新南威尔士大学)
;
Wake Forest University(维克森林大学)
Detection of adversarial intent in Human-AI teams using LLMs
利用大语言模型检测人类-人工智能团队中的对抗意图
Abed K. Musaffar, Ambuj Singh, Francesco Bullo
机构
*
Department of Mechanical Engineering, University of California at Santa Barbara(加州大学圣巴巴拉分校机械工程系)
;
Department of Computer Science, University of California at Santa Barbara(加州大学圣巴巴拉分校计算机科学系)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
对可信监控的自适应攻击颠覆AI控制协议
Mikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre, Maksym Andriushchenko, Ameya Prabhu, Jonas Geiping
机构
*
MATS
;
EPFL(瑞士联邦理工学院)
;
ELLIS Institute Tübingen & Max Planck Institute for Intelligent Systems(图宾根ELLIS研究所及图宾根马克斯·普朗克智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
;
University of Tübingen(图宾根大学)