Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
机构 * Department of Mathematics, UCLA(UCLA数学系) ; Department of Mathematical Sciences, Seoul National University(首尔国立大学数学科学系) ; KRAFTON ; Department of Linguistics, Stanford University(斯坦福大学语言学系) ; Department of Computer Science and Engineering, Santa Clara University(圣克拉拉大学计算机科学与工程系)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG