Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
反馈归一化的开发者内存用于强化学习编码代理:一种安全门控MCP架构
机构 * PythaLab, Yildiz Technical University, Istanbul, Türkiye(PythaLab,伊兹密尔技术大学,伊斯坦布尔,土耳其)
专题命中 RAG评测 :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL
AI总结 本文提出RL Developer Memory架构,通过归一化反馈和安全门控机制提升强化学习编码代理的内存管理,实验证明其在确定性任务中的有效性。
Comments 25 pages, 5 figures, 7 tables. Preprint. Implementation and supplementary artifacts are available at the project repository