Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making
混合序列建模与强化验证的可控目标条件决策
机构 * School of Artificial Intelligence, Beihang University(北航人工智能学院) ; Department of Computing Science and Amii, University of Alberta(阿尔伯塔大学计算机科学系和Amii) ; Edmonton Research Center, Huawei Canada(华为加拿大埃德蒙顿研究中心) ; Hangzhou International Innovation Institute, Beihang University(杭州国际创新院,北航) ; School of Artificial Intelligence and Computer Science, North China University of Technology(北中国技术大学人工智能与计算机科学学院) ; Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory(鹏城实验室多智能体与具身智能研究院)
AI总结 提出Doctor框架,结合掩码轨迹Transformer与强化验证,通过生成候选动作并选择验证值最接近目标回报的动作,提升目标条件策略在低覆盖数据下的可控性。