Cognitive models can reveal interpretable value trade-offs in language models
认知模型可以揭示语言模型中的可解释价值权衡
机构 * Kempner Institute for Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究所) ; Google DeepMind(谷歌DeepMind) ; Department of Psychology, Harvard University(哈佛大学心理学系)
专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本文提出利用认知模型评估语言模型中的价值权衡,揭示其行为特征变化及对社会行为的影响。
Comments 10 pages, 5 figures