Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
对齐、正交或冲突:何时可以安全地优化思维链?
机构 * Google DeepMind(谷歌DeepMind)
专题命中 其他推理 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.AI、cs.LG
AI总结 本文探讨了在训练过程中思维链(CoT)的可监控性受训练影响的问题,提出了一种框架来预测何时以及为何发生冲突,并通过实验验证了正交和冲突奖励项对CoT监控性的影响。