Monitoring Emergent Reward Hacking During Generation via Internal Activations
通过内部激活监控生成过程中的涌现奖励黑客行为
机构 * Technical University of Berlin(柏林技术大学) ; BIFOLD - Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);test-time compute(abstract);分类 cs.CL、cs.AI
AI总结 通过分析内部激活来监控生成过程中的奖励黑客行为,提升微调语言模型的安全性。
Journal ref ICLR2026 Workshop: Principled Design for Trustworthy AI