Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
专题命中 后训练与偏好优化 :LLM(title);language model(abstract);分类 cs.CL、cs.LG
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
专题命中 后训练与偏好优化 :LLM(title);language model(abstract);分类 cs.CL、cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments ICML2025
机构 * City University of Hong Kong(香港城市大学) ; Zhejiang University(浙江大学) ; KTH Royal Institute of Technology(皇家理工学院) ; Hangzhou Dianzi University(杭州电子科技大学)
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
机构 * University of Technology Sydney(技术科技大学) ; University of Washington(华盛顿大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Zhejiang University(浙江大学)
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI、cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments 36 Pages; ICML 2025
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI、cs.LG
Comments Accepted to CVPR 2025 as a highlight paper
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.AI、cs.LG
Comments 8 pages
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.LG
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments Updated for AAMAS 2025 camera-ready. This preprint represents the full version of the paper, including all proofs, experimental details, and additional discussions
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL、cs.LG
Comments 20 pages, 12 figures
专题命中 后训练与偏好优化 :LLM(title,abstract);分类 cs.CL、cs.AI
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.AI、cs.LG
Comments Accepted by AAAI 2025
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.AI
Comments EMNLP 2024
Journal ref Proc. EMNLP (2024) 9004-9018
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.AI、cs.LG
Comments 22 pages, 7 figures, 4 tables
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments 64 pages. 14 pages for main paper, 50 pages for references + appendix
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL、cs.LG
Comments To appear in NeurIPS 2024 in the Fine-Tuning in Machine Learning Workshop
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL、cs.AI
Comments EMNLP 2024 main conference
专题命中 后训练与偏好优化 :post-training(title,abstract);分类 cs.AI、cs.LG
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL、cs.LG
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.CL、cs.AI
Comments Accepted at ACM MM 2024
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.AI
Comments link: https://hf.co/spaces/WildVision/vision-arena
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.LG
Comments 8 pages, 4 figures
专题命中 后训练与偏好优化 :LLM(title,abstract);分类 cs.AI、cs.LG
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.CL、cs.AI
Comments ICLR 2024
专题命中 后训练与偏好优化 :language model(title,abstract);分类 cs.AI、cs.LG
Comments Presented at International Conference on Learning Representations (ICLR) 2024
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.CL、cs.LG
专题命中 后训练与偏好优化 :language agent(title,abstract);分类 cs.CL、cs.AI
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.AI、cs.LG
Comments Accepted for presentation at NeurIPS 2023; 29 pages
专题命中 后训练与偏好优化 :RLHF(title,abstract);分类 cs.CL、cs.LG