Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations
多买家市场中的战略谈判:基于可验证奖励的强化学习用于大语言模型谈判
Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Xin Chen, David Simchi-Levi
机构
*
Institute for Data, Systems, and Society, Massachusetts Institute of Technology(数据、系统与社会研究所,麻省理工学院)
;
Mitch Daniels School of Business, Purdue University(米奇·丹尼尔斯商学院,普渡大学)
;
The Division of Physics, Mathematics and Astronomy, California Institute of Technology(物理、数学与天文学部,加州理工学院)
;
Department of Computing and Mathematical Sciences, California Institute of Technology(计算与数学科学系,加州理工学院)
;
H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology(H. 米尔顿·斯图尔特工业与系统工程学院,佐治亚理工学院)
;
Department of Civil and Environmental Engineering, Operations Research Center, Massachusetts Institute of Technology(土木与环境工程系、运筹学中心,麻省理工学院)
专题命中
后训练与偏好优化
:LLM(title);large language model(abstract);language model(abstract);分类 cs.LG
LaGO: Latent Action Guidance for Online Reinforcement Learning
LaGO:面向在线强化学习的潜在动作引导
Kuan-Yen Liu, Ren-Jyun Huang, Ti-Rong Wu
机构
*
Siebel School of Computing(计算科学系)
;
Data Science, University of Illinois Urbana-Champaign, USA(数据科学,伊利诺伊大学厄巴纳-香槟分校,美国)
;
Department of Computer Science, National Yang Ming Chiao Tung University, Taiwan(计算机科学系,National Yang Ming Chiao Tung大学,台湾)
;
Institute of Information Science, Academia Sinica, Taiwan(信息科学研究所, Academia Sinica,台湾)
专题命中
后训练与偏好优化
:LLM(abstract,abstract_cn);large language model(abstract,comments);language model(abstract,comments);分类 cs.AI
Comments8 pages (21 in total), 2 figures, 4 tables. Original manuscript for a Guided Research project conducted at the Technical University of Munich, detailing the complete methodology, full data pipeline, and comprehensive experimental results. A related, condensed subset of this work was subsequently adapted and published at the StereACuLT 2026 workshop
机构
*
Generative AI Lab, College of Computing and Data Science, Nanyang Technological University(生成式人工智能实验室,南洋理工大学计算与数据科学学院)
;
Tongyi Lab, Alibaba Group(阿里集团通义实验室)
专题命中
后训练与偏好优化
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Active-GRPO:用于分子优化的自适应模仿与自我改进推理
Xuefeng Liu, Mingxuan Cao, Qinan Huang, Thomas Brettin, Rick Stevens, Le Cong
机构
*
School of Medicine, Stanford University(斯坦福大学医学院)
;
Data Science Institute, University of Chicago(芝加哥大学数据科学研究所)
;
Pritzker School of Molecular Engineering, University of Chicago(芝加哥大学普利兹克分子工程学院)
;
Department of Computer Science, University of Chicago(芝加哥大学计算机科学系)
;
Argonne National Laboratory(阿贡国家实验室)
专题命中
后训练与偏好优化
:SFT(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Comments28 pages. When AI learns from human feedback, it forces a single "correct" answer, but sometimes multiple answers are all genuinely valid, and that nuance gets thrown away