SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing
SLIM-RL: 无需轨迹切分的扩散LLM风险预算随机掩码强化学习
Ruikang Zhao, Zhenting Wang, Han Gao, Ligong Han
机构
*
Technical University of Denmark(丹麦技术大学)
;
MBZUAI Institute of Foundation Models(穆罕默德·本·扎耶德人工智能大学基础模型研究所)
;
Iowa State University(爱荷华州立大学)
;
Red Hat AI Innovation(红帽人工智能创新实验室)
;
MIT–IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室)
机构
*
Hong Kong University of Science and Technology(香港科技大学)
;
Tencent VISVISE(腾讯VISVISE)
;
Peking University(北京大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Texas A&M University(德克萨斯大学)
;
Macau University of Science and Technology(澳门科技大学)
Goal-oriented learning of stochastic differential equations using error bounds on path-space observables
基于路径空间观测误差界限的随机微分方程目标导向学习
Joanna Zou, Han Cheng Lie, Youssef Marzouk
机构
*
Center for Computational Science & Engineering, Massachusetts Institute of Technology(麻省理工学院计算科学与工程中心)
;
Department of Mathematics, University of Potsdam(波茨坦大学数学系)