A Descriptive and Normative Theory of Human Beliefs in RLHF
基于人类反馈强化学习中的人类信念描述性与规范性理论
机构 * College of Information and Computer Sciences University of Massachusetts Amherst(信息与计算机科学学院 马萨诸塞大学阿姆赫斯特分校) ; Department of Computer Science The University of Texas at Austin(计算机科学系 德州大学奥斯汀分校)
专题命中 后训练与偏好优化 :RLHF(title,summary_cn);分类 cs.AI、cs.LG
AI总结 研究 RLHF 中人类信念作用,提出新偏好模型,通过人类和综合实验,证实人类对智能体能力信念影响偏好,指出减少信念与能力不匹配可提升 RLHF 并给出新实践方向。
Comments Published at TMLR