The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
规范陷阱:为何仅靠静态价值对齐无法实现稳健对齐
机构 * Belmont University(贝尔蒙特大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract,abstract_cn);分类 cs.AI、cs.CY、cs.LG
AI总结 本文指出静态内容导向的人工智能价值对齐在能力扩展、分布偏移和自主性提升时无法实现稳健对齐,探讨了哲学难题及现有方法的结构性漏洞。
Comments 31 pages, no figures. Version 5. First posted as arXiv:2512.03048 in November 2025. First in a six-paper research program on AI alignment