Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
隐秘动机:在连续思维模型中检测不一致的推理
机构 * Stanford University(斯坦福大学)
专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)
AI总结 本文研究了连续思维模型中如何检测不一致的推理,提出MoralChain基准测试,并通过双触发方法训练模型,揭示了潜在空间中不一致推理的存在及早期规划阶段的安全监控重要性。
Comments 15 pages with 2 figures
Journal ref International Conference on Learning Representations (ICLR) Latent & Implicit Thinking (LIT) Workshop 2026