OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
OPID: 面向智能体强化学习的在策略技能蒸馏
Shuo Yang, Jinyang Wu, Zhengxi Lu, Yuhao Shen, Fan Zhang, Lang Feng, Shuai Zhang, Haoran Luo, Zheng Lian, Zhengqi Wen, Jianhua Tao
机构
*
Tsinghua University(清华大学)
;
Zhejiang University(浙江大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Nanyang Technological University(南洋理工大学)
;
Tongji University(同济大学)
Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations
安全关键场景下基于蒸馏视觉语言模型的事件自适应运动规划
Zhenwei Huang, Changsheng You, Shuai Wang, Chao Zhou, Wei Xu, Yi Gong
机构
*
Southern University of Science and Technology(南方科技大学)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Manifold Tech Limited(曼孚科技股份有限公司)
机构
*
the Siebel School of Computing and Data Science(塞比尔计算与数据科学学院)
;
the Department of Agricultural and Biological Engineering(农业与生物工程系)
;
National Center for Supercomputing Applications(国家超级计算中心)
A Generalization Theory for JEPA-Based World Models
基于JEPA的世界模型的泛化理论
Jingyi Cui, Qi Zhang, Hongwei Wen, Yisen Wang
机构
*
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室)
;
University of Sydney(悉尼大学)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Department of Networked Intelligence, Peng Cheng Laboratory(鹏城实验室网络智能部)
;
School of Computer Science and Engineering, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)计算机科学与工程学院)