Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
嵌入空间中的隐秘行动:基于几何的隐写术与大语言模型中的检测
机构 * UCL Centre for Artificial Intelligence, University College London, UK(伦敦大学学院人工智能中心,大学学院伦敦) ; School of Computing and Communications, Lancaster University, UK(兰卡斯特大学计算机与通讯学院) ; Department of Informatics, University of Sussex, UK(苏塞克斯大学信息学院) ; Centre for Psychedelic Research and Centre for Complexity Science, Imperial College London, UK(伦敦帝国学院迷幻研究与复杂科学中心) ; Centre for Eudaimonia and Human Flourishing, University of Oxford, UK(牛津大学幸福与人类繁荣中心)
AI总结 本研究提出了一种低恢复性隐写术,通过嵌入空间衍生映射提升秘密恢复率,同时减少负载恢复性,并提出基于机制可解释性的检测方法,提高微调模型中的秘密检测准确率。