The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI
实验室驱动的对齐签名的出现:一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险
机构 * AI Researcher(人工智能研究员)
专题命中 AI治理与伦理 :alignment(title);分类 cs.CL
AI总结 本文提出了一种心理测量框架,用于审计生成AI中的潜在偏见和叠加风险,通过分析九个领先模型的实验室信号,揭示了持续行为聚类的成因。
Comments v2: expanded from 9 to 18 behavioral dimensions and from 4 to 6 developer organizations; revised statistical methodology (rank-based inference with effect-size criterion, replacing variance-decomposition approach); model-level results now reported; references corrected throughout