The Polite Liar: Epistemic Pathology in Language Models
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 17 pages, 2 tables, Preprint - under review at AI & Society
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 17 pages, 2 tables, Preprint - under review at AI & Society
机构 * Rensselaer Polytechnic Institute(罗切斯特理工学院) ; IBM Research(IBM研究院) ; Cornell University(康奈尔大学)
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * National Centre for Text Mining, The University of Manchester(曼彻斯特大学文本挖掘中心) ; School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) ; Center for Language and Information Research, Wuhan University(武汉大学语言与信息研究中心) ; The Fin AI(Fin AI公司) ; Baidu Inc.(百度公司) ; Archimedes Research(阿基米德研究)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by the EMNLP 2025 main conference
Journal ref https://aclanthology.org/2025.emnlp-main.359/
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025
机构 * Yale University(耶鲁大学) ; Allen Institute for AI(人工智能研究院)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Computer Science, McGill University(麦吉尔大学计算机科学系) ; Mila - Quebec AI Institute(魁北克人工智能研究所) ; UC Berkeley(加州大学伯克利分校) ; G o o g l e , Paradigms of Intelligence Team(谷歌,智能范式团队) ; Mathematics and Statistics, Université de Montréal(蒙特利尔大学数学与统计学系) ; Neurology & Neurosurgery and Montreal Neurological Institute, McGill University(神经病学与神经外科及蒙特利尔神经研究所,麦吉尔大学) ; CIFAR Learning in Machines & Brains Program(CIFAR 机器与大脑学习计划)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 33 pages, 14 figures, 9 tables
机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; Cambium Assessment(Cambium评估)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.CY、cs.LG
Comments Published in EMNLP 2025: The 2025 Conference on Empirical Methods in Natural Language Processing
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Project page, code & data: https://machine-bullshit.github.io
机构 * Peking University(北京大学) ; WeChat AI(微信AI) ; William & Mary(威廉与玛丽学院) ; Westlake University(西交利物浦大学)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 25 pages, 9 figures, Code & model weights available at: https://zhuohaoyu.github.io/RewardAnything
机构 * SAGEA ; Tribhuwan University(特里布文大学) ; Vedas College(维达学院) ; Madan Bhandari Memorial College(马达恩·班达里纪念学院) ; Fudan University(复旦大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 偏好对齐 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 19 pages, 2 figures, 9 tables
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) ; University of Illinois Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校) ; Korea University(韩国大学)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICML 2025
机构 * Independent Researcher(独立研究者) ; Machine Intelligence Research Institute(机器智能研究院)
专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Preprint. This work has been submitted to the Reliable and Responsible Foundation Models Workshop at ICML 2025 for review
专题命中 偏好对齐 :RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Centre for Machine Intelligence and Data Science, IIT Bombay(印度理工学院班加罗尔机器智能与数据科学中心) ; Koita Centre for Digital Health, IIT Bombay(印度理工学院班加罗尔数字健康中心) ; Department of Computer Science and Engineering, IIT Bombay(印度理工学院班加罗尔计算机科学与工程系)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Updated abstract, algorithm and experimental results
机构 * FAIR at Meta(Meta 的 FAIR 部门) ; University of Washington(华盛顿大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments CVPR 2025 Camera-ready. Project page: https://silmm.github.io/
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 15 pages, 3 figures, 7 tables
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 22 pages, 10 figures
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 22 pages, 16 figures, 7 tables
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at NeurIPS 2024
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by COLING2025
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2024