Improving Socratic Question Generation using Data Augmentation and Preference Optimization
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.CY、cs.LG
Comments Published at the 19th BEA Workshop co-located with NAACL-2024
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.CY、cs.LG
Comments Published at the 19th BEA Workshop co-located with NAACL-2024
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NAACL 2024
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Models and data are available at https://github.com/OpenBMB/Eurus
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NAACL2024 Main Track Long Paper
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 8 pages
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2024. Code is available on our project website: https://xingyaoww.github.io/mint-bench
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments First two authors contributed equally; Project website: https://selma-t2i.github.io/
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, 9 figures
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Updates for camera-ready submission
Journal ref NeurIPS Workshop on Generative AI for Education (GAIED), 2023
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 25 pages, 6 figures
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2023 main conference
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments technical report. arXiv admin note: text overlap with arXiv:2306.16636 by other authors
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Proceedings of the 40th International Conference on Machine Learning (ICML), 2023
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint. Code at https://github.com/FranxYao/chain-of-thought-hub
专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments for associated data visualizations, see https://www.evals.anthropic.com/model-written/ for full datasets, see https://github.com/anthropics/evals
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments AIES 2020
可引导的文化偏好优化奖励模型
机构 * Stanford University(斯坦福大学) ; University of Amsterdam(阿姆斯特丹大学)
专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.CL、cs.AI
AI总结 提出SCPO算法,通过平衡多种文化偏好训练奖励模型,在PRISM和GlobalOpinionQA数据集上提升少数群体偏好预测准确率最多7点,训练效率提高280%。
Comments Accepted to Pluralistic Alignment @ ICML 2026
机构 * Duke University(杜克大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.AI、cs.CY
Comments To appear in the AAAI 2026 Alignment Track
机构 * KAIST(韩国科学技术院) ; Columbia University(哥伦比亚大学)
专题命中 偏好对齐 :RLHF(abstract,comments);分类 cs.AI、cs.LG
Comments Accepted at ACL 2025, Source code: https://github.com/mintaywon/IF_RLHF
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 63 (2025) 27471-27500
机构 * Department of Electrical and Computer Engineering, University of Texas at San Antonio, Texas, USA(电子与计算机工程系,德克萨斯州立大学圣安东尼奥分校) ; DEVCOM Army Research Lab, USA(陆军研究实验室)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG;alignment(comments)
Comments Accepted to the workshop on Models of Human Feedback for AI Alignment at the 42nd International Conference on Machine Learning
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG;trustworthy(comments)
Comments Accepted in TrustNLP: Third Workshop on Trustworthy Natural Language Processing, co-located with ACL 2023