Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL
生存或崩溃:自我博弈强化学习中数据门控与奖励基础的不对称作用
机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) ; Cisco Research(思科研究)
AI总结 本文研究了自我博弈强化学习中数据门控和奖励基础的不对称作用,发现数据门控是维持稳定的关键因素,而奖励信号在门控移除后无法单独保证稳定性,揭示了'基础提出者悖论'。