Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
受机器人启发的用于社会敏感领域基础模型的护栏
机构 * Yale University(耶鲁大学) ; Kyoto University(京都大学)
AI总结 本文提出了一种基于机器人学的护栏框架,用于在社会敏感领域中对基础模型进行运行时行为控制,以减少交互轨迹中向不良状态的漂移,并适应多样化的社会情境。
Comments Under review at Journal of Artificial Intelligence Research (JAIR)