Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
基于熵的大型语言模型价值漂移与对齐工作的测量
专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出基于熵的框架,用于测量大型语言模型中的价值漂移和对齐工作,通过五种行为分类法和熵动态分析,实现对模型伦理熵的实时监控与警报。
Comments 6 pages. Companion paper to "The Second Law of Intelligence: Controlling Ethical Entropy in Autonomous Systems". Code and tools: https://github.com/AerisSpace/EthicalEntropyKit