Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
将LLM后训练为更好的决策智能体:一种遗憾最小化方法
机构 * Massachusetts Institute of Technology(麻省理工学院) ; University of Maryland, College Park(马里兰大学哥伦比亚学院)
专题命中 后训练与偏好优化 :LLM(title_cn,summary_cn);post-training(title,abstract);large language model(abstract);language model(abstract)
AI总结 提出迭代遗憾最小化微调(Iterative RMFT),通过反复蒸馏低遗憾决策轨迹来后训练LLM,提升其在在线决策任务中的表现,无需依赖已知算法或人工模板。
Comments Camera ready version of ICML 2026