Multi-Objective Exploration and Preference Optimization via Mutual Information
基于互信息的多目标探索与偏好优化
机构 * School of Computer, Beihang University(北京航空航天大学计算机学院) ; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd(星辰AGI实验室,中国电信人工智能技术(北京)有限公司)
专题命中 后训练与偏好优化 :preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 提出MI-EPO框架,通过最大化生成响应、偏好反馈和偏好向量的联合条件互信息,统一多目标探索与对齐,实现可控且稳定的多目标权衡。
Comments Accepted at ECML/PKDD 2026