arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-04 至 2026-05-04 共收录 1 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 1 篇

2605.00224 2026-05-04 cs.AI 92%

TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

TUR-DPO:基于拓扑和不确定性的直接偏好优化

Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad Oussalah

机构 * Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(人工智能与创新中心,乌尔米耶大学,伊拉克) Department of Computer Engineering, University of Kurdistan, Iran(计算机工程系,乌尔米耶大学,伊朗) Centre for Artificial Intelligence Research and Optimisation, Torrens University Australia, Brisbane, Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚,布里斯班,澳大利亚) Research and Innovation Center, Obuda University, Budapest 1034, Hungary(研究与创新中心,奥布达大学,布达佩斯1034,匈牙利) Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland(机器视觉与信号分析中心(CMVS),奥卢大学,芬兰)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(abstract,abstract_cn);分类 cs.AI

AI总结 TUR-DPO通过引入轻量级推理拓扑和结合语义忠实度、效用和拓扑质量,提升偏好对齐的稳定性与鲁棒性,同时保持训练简洁性和无需在线回滚。

Comments Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏