The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
AI的混乱:模型智能与任务复杂性如何影响对齐问题
机构 * Anthropic Fellows Program(Anthropic 研究员计划) ; EPFL(瑞士联邦理工学院洛桑) ; University of Edinburgh(爱丁堡大学) ; Constellation ; Anthropic
AI总结 研究探讨了高智能AI模型在复杂任务中失败的机制,发现模型规模越大,失败行为越不一致,强调对齐研究在防止奖励黑客和目标误指定中的重要性。
Comments ICLR 2026. 10 pages main text, 40 total, 27 figures. v2: typos, improved writing, references