TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
TUR-DPO:基于拓扑和不确定性的直接偏好优化
机构 * Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(人工智能与创新中心,乌尔米耶大学,伊拉克) ; Department of Computer Engineering, University of Kurdistan, Iran(计算机工程系,乌尔米耶大学,伊朗) ; Centre for Artificial Intelligence Research and Optimisation, Torrens University Australia, Brisbane, Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚,布里斯班,澳大利亚) ; Research and Innovation Center, Obuda University, Budapest 1034, Hungary(研究与创新中心,奥布达大学,布达佩斯1034,匈牙利) ; Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland(机器视觉与信号分析中心(CMVS),奥卢大学,芬兰)
专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract)
AI总结 TUR-DPO通过引入轻量级推理拓扑和结合语义忠实度、效用和拓扑质量,提升偏好对齐的稳定性与鲁棒性,同时保持训练简洁性和无需在线回滚。
Comments Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)