Fragility of Value under Imperfect Alignment
不完美对齐下价值的脆弱性
专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文针对AI系统与人类价值对齐问题,建立模型明确人类价值函数与代理条件准确度的相关条件,凸显过度优化风险,提出采用quantilizers等限制优化压力的AI设计方案。
Comments 25 pages, 7 figures. Expanded the contribution statements and added three researchers as coauthors. Made small text improvements