Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
迈向用于小规模语言模型智能体的稳健强化学习
机构 * University of Waterloo(滑铁卢大学) ; Khulna University of Engineering & Technology(库尔纳工程技术大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究小规模语言模型强化学习不稳定问题,识别出三种失败模式,提出容量余量假设,采用合并并重新初始化适配器技术等方法,所提系统稳定收敛,提升偏好胜率,优于指令调整基线且减少训练数据。
Comments Proceedings of the 2026 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Bellevue, WA, USA