SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
SafeHarbor:用于LLM智能体安全的分层记忆增强防护栏
机构 * School of Cyber Science and Technology, Beihang University, Beijing, China(北京航空航天大学网络安全学院) ; Institute of Artificial Intelligence, Beihang University, Beijing, China(北京航空航天大学人工智能研究院) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) ; AI Security Lab, Beijing, China(360人工智能安全实验室)
AI总结 提出SafeHarbor框架,通过分层记忆系统和信息熵自进化机制,在保持高安全拒绝率的同时提升对良性请求的响应能力。
Comments Accepted by ICML 2026