Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow Mechanism
非对抗模仿学习无复合误差的理论保证:价值流机制
机构 * National Key Laboratory for Novel Software Technology(新型软件技术国家实验室) ; School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Shenzhen Research Institute of Big Data(深圳大数据研究院)
AI总结 本文证明IQ-Learn等价于行为克隆,存在复合误差,并提出基于Bellman约束的Dual Q-DM方法,通过价值流传播实现非对抗模仿学习,理论上消除复合误差。
Comments Paper accepted by ICML 2026