Learning When to Trust via Selective Context Preference Optimization
通过选择性上下文偏好优化学习何时信任
机构 * Duke University(杜克大学) ; National University of Singapore(新加坡国立大学) ; UC Berkeley(加州大学伯克利分校) ; UC Irvine(加州大学欧文分校) ; Northeastern University(东北大学) ; Nanyang Technological University(南洋理工大学)
AI总结 本研究针对语言模型易受误导性外部信号影响的问题,提出SCOPE方法,在含四种条件的MIST基准上优化DPO目标,降低SC2W的同时保持正常上下文下的准确率,主张以选择性信任评判模型。
Comments Project Page at this https URL (https://worldbench.github.io/scope) GitHub Repo at this https URL (https://github.com/worldbench/SCOPE) HF Dataset at this https URL (https://huggingface.co/datasets/worldbench/MIST-Bench)