A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
多语言语言模型毒性检测与缓解策略综述
机构 * Scale AI ; Indian Institute of Technology Gandhinagar(印度理工学院甘地讷格尔分校) ; University of Virginia(弗吉尼亚大学)
AI总结 综述多语言大模型的毒性检测与缓解方法,包括威胁模型、检测方法(跨语言编码器、翻译流水线等)和缓解策略(数据过滤、偏好调优等),指出语言覆盖不均、文化依赖等挑战。
Comments Accepted to the Findings of ACL, 2026
Journal ref Findings of ACL, 2026