Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
具有新鲜度意识的优先经验回放用于LLM/VLM强化学习
机构 * King Abdullah University of Science and Technology(卡布尔大学科学与技术大学) ; Chinese Academy of Sciences, Institute of Automation(中国科学院自动化研究所) ; AI Centre, Department of Computer Science, University College London(伦敦大学学院人工智能中心) ; Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究所)
专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)
AI总结 本文提出Freshness-Aware PER,通过引入指数衰减机制解决LLM/VLM强化学习中优先级过时问题,显著提升样本效率,在多个任务中取得优异表现。