arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

NeurIPS

Conference on Neural Information Processing Systems · 会议 · Machine Learning

2026-06-24 至 2026-06-24 共收录 1
2509.03647 2026-06-24 cs.CL cs.AI cs.LG 版本更新

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

打破镜像:基于激活的LLM评估者自我偏好缓解方法

Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer

机构 * University of Virginia(弗吉尼亚大学) University of California, San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学) School of Computer Science(计算机科学学院)

AI总结 针对LLM评估者自我偏好偏见,提出轻量级引导向量方法,在推理时无需重训练即可将不公正自我偏好降低97%,但存在稳定性问题。

Comments Presented at {Mechanistic Interpretability, Evaluations, Reliable-ML} Workshops, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏