Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
尾感知信息论泛化用于RLHF和SGLD
机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) ; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高精尖创新中心) ; Advanced Institute of Information Technology, Peking University(北京大学信息技术高等研究院) ; Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所) ; Computer and Mathematical Sciences, Computer Science, and Statistics, University of Toronto(多伦多大学计算机与数学科学、计算机科学和统计学系) ; MBZUAI(穆罕默德·本·扎耶德人工智能大学)
专题命中 偏好对齐 :RLHF(title_cn,abstract);alignment(title,abstract);分类 cs.AI、cs.LG
AI总结 本文提出一种尾依赖信息论框架,用于处理亚韦尔bull数据,通过移位对数f_θ-分歧来界定制换测度期望,从而在重尾数据下获得更精确的泛化界限。
Comments 47 pages, 6 figures