Comments17 pages, 12 figures, 14 tables. Preprint; under review at EACL 2027 (ACL Rolling Review, August 2026 cycle). Code and data: https://github.com/mohammadi-hadi/MAP-PO
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Xiaomi Corporation(小米公司)
;
State Key Laboratory of Communication Content Cognition, People’s Daily Online(人民日报社传播内容认知国家重点实验室)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
Vellore Institute of Technology(维洛雷理工学院)
;
University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)
;
Northwestern University(西北大学)
;
Yale University(耶鲁大学)
;
Algoverse AI Research(Algoverse AI研究)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Comments12 pages, 4 figures, 6 tables. Includes ablation study across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 5 math reasoning benchmarks (GSM8K, MATH500, Minerva, AIME24, Gaokao2023). GPT-4.1 used for structured evaluation of reasoning quality
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
NJIT(新泽西理工学院)
;
Jilin University(吉林大学)
;
Institute of Computing Technology, CAS(中国科学院计算技术研究所)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
后训练与偏好优化
:preference optimization(title);LLM(abstract_cn);large language model(abstract);language model(abstract)
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
偏好优化中的虚假相关学习:机制、后果及通过平局训练的缓解方法
Christian Moya, Alex Semendinger, Guang Lin, Elliott Thornley
机构
*
Department of Mathematics, Purdue University, West Lafayette IN, USA(普渡大学数学系)
;
School of Mechanical Engineering, Purdue University, West Lafayette IN, USA(普渡大学机械工程学院)
;
Massachusetts Institute of Technology, Cambridge MA, USA(麻省理工学院)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
机构
*
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR
在线强化学习需要多少?用于RLVR中离线偏好优化的信息性回放
Richa Verma, Balaraman Ravindran
机构
*
TCS Research Department of CSE(TCS计算机科学系研究部)
;
IIT Madras(印度理工学院马德拉斯分校)
;
Department of Data Science & AI(数据科学与人工智能系)
;
Wadhwani School of Data Science & AI(Wadhwani数据科学与人工智能学院)