Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
道德敏感性与在角色扮演下大型语言模型的鲁棒性
机构 * TELUS Digital Research Hub(TELUS数字研究中心) ; Center for Artificial Intelligence and Machine Learning(人工智能与机器学习中心) ; Institute of Mathematics, Statistics and Computer Science(数学、统计与计算机科学研究所) ; University of São Paulo(圣保罗大学)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);post-training(abstract)
AI总结 本文研究了大型语言模型在角色扮演下的道德反应,通过道德基础问卷基准测试,量化了道德敏感性和鲁棒性,揭示了模型家族对鲁棒性的影响显著,而预训练对敏感性起主导作用。
Comments Added experiments with a logit-based method and now reporting unbounded metrics