Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
无最优演示者的逆强化学习:一种可行奖励集方法
Kihyun Kim, Shripad Deshmukh, Nikos Vlassis, Jiawei Zhang
机构
*
MIT LIDS(麻省理工学院媒体实验室)
;
University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校)
;
Adobe Research(Adobe研究院)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
学习的中继表示用于前向思考的离散扩散模型
Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum
机构
*
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
;
University of Toronto(多伦多大学)
;
Stanford University(斯坦福大学)
;
Imperial College London(伦敦帝国学院)
;
Mila
;
Vijil