Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
机构 * Georgia Institute of Technology(佐治亚理工学院) ; University of Bristol(布里斯托大学) ; University of Maryland(马里兰大学) ; University College London(伦敦大学学院) ; MATS ; Astra ; MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中 知识编辑与模型理解 :LLM(abstract,comments);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Code at https://github.com/aengusl/latent-adversarial-training. Models at https://huggingface.co/LLM-LAT