Hierarchical Alignment: Surgical Fine-Tuning via Functional Layer Specialization in Large Language Models
机构 * The Chinese University of Hong Kong(香港中文大学) ; Fudan University(复旦大学)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * The Chinese University of Hong Kong(香港中文大学) ; Fudan University(复旦大学)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
机构 * University of Southern California(南加州大学) ; California State University(加州州立大学)
专题命中 偏好对齐 :DPO(title,abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) ; RUC(中国人民大学) ; USTC(University of Science and Technology of China) ; NTU(National University of Technology) ; NUS(National University of Singapore) ; USC(University of Southern California) ; SCU(Sichuan University) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.LG
机构 * Department of Electrical & Computer Engineering(电气与计算机工程系) ; Boston University(波士顿大学) ; Department of Computer Science(计算机科学系) ; Cornell University(康奈尔大学) ; Faculty of Computing & Data Sciences(计算与数据科学学院)
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
机构 * University of Washington(华盛顿大学) ; Xi’an Jiaotong University(西安交通大学) ; University of Notre Dame(圣母大学) ; Google(谷歌)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.LG
Comments 8 pages, 1 figures, 2 tables. Experimental code and results are publicly available at https://anonymous.4open.science/r/Graph_RL-BF08/readme.md
机构 * Department of Computer Science(计算机科学系) ; Purdue University(普渡大学)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
Comments Accepted by COLM 2025
机构 * Department of Computer Science(计算机科学系)
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * The University of Texas at Austin, US(德克萨斯大学奥斯汀分校)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Google Cloud AI Research(谷歌云人工智能研究) ; Google DeepMind(谷歌DeepMind) ; University of Virginia(弗吉尼亚大学)
专题命中 偏好对齐 :alignment(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI
Comments 26 pages
机构 * Department of Statistics, Rutgers University, New Brunswick, United States ; Department of Computer Science, Rutgers University, New Brunswick, United States ; College of Management of Technology, EPFL, Switzerland ; Department of Computer Science, ETH Zurich, Switzerland
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments ICML 2025
机构 * Peking University(北京大学)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
机构 * Institute for Computational and Mathematical Engineering (ICME), Stanford University(计算与数学工程研究所(ICME),斯坦福大学) ; Amazon(亚马逊) ; Department of Mathematics, Stanford University(数学系,斯坦福大学)
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
Comments Published at UAI 2025
机构 * University of Oxford(牛津大学) ; Jagiellonian University(雅盖隆大学) ; Harvard University(哈佛大学)
专题命中 偏好对齐 :DPO(title,abstract);safety(abstract);分类 cs.CL、cs.LG
Journal ref NeurIPS 2024 Workshop on Socially Responsible Language Modelling Research (SoLaR)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments Camera ready version for ACL 2025 Findings
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI
Comments Accepted at ICML 2025
专题命中 偏好对齐 :alignment(title);RLHF(abstract);harmlessness(abstract);分类 cs.AI、cs.LG
Comments Accepted at ACL2025 Findings
机构 * Meta ; University of Chicago(芝加哥大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
Comments Code and dataset are available at https://teachingwithlies.github.io/
机构 * Northwestern University(西北大学) ; RIKEN-AIP(日本理化学研究所-AIP) ; The University of Tokyo(东京大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments Language Modeling, Machine Learning for NLP, Distributional Pareto-Optimal
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG
Comments Published at ICML 2025
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments spotlight @ neurips language gamification workshop. updated the problem description and added new online RL experiments in this version
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments Accepted to ICLR 2025
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments This paper has been accepted to ICLR 2025
专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI、cs.LG