Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation
合成轨迹反映真实的奖励黑客行为吗?对监控真实世界代码生成中黑客行为的系统研究
机构 * Peking University(北京大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; Arena ; University of Maryland(马里兰大学) ; Mohamed bin Zayed University of Artificial Intelligence(莫莫迪 bin Zayed 人工智能大学)
专题命中 代码生成 :code generation(title,abstract);分类 cs.LG
AI总结 本文研究合成轨迹与真实世界中代码生成奖励黑客行为的差异,通过对比监控器在合成与真实数据上的表现,发现合成数据训练的监控器无法泛化到真实黑客行为,而真实数据训练的监控器对未见过的黑客类型更具泛化能力。
Comments Corrected a typo in the title; no other changes