Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning
基于组合子目标评分的方向条件策略用于在线目标条件强化学习
Swaminathan S K, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur(计算机科学与工程系,印度理工学院Kharagpur分校)
;
Department of Mechanical Engineering, Indian Institute of Technology Kharagpur(机械工程系,印度理工学院Kharagpur分校)
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
QPILOTS:面向流策略的高效测试时Q引导
Yifan Ruan, Chenyang Cao, Andreas Burger, Ali Pesaranghader, Kaveh Kamali, Jaehong Kim, Nandita Vijaykumar, Alan Aspuru-Guzik, Igor Gilitschenski, Nicholas Rhinehart
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
LG Electronics(LG电子)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
TimeRewarder: 通过帧间时间距离从被动视频中学习密集奖励
Yuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman, Yang Gao
机构
*
Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, China(清华大学交叉信息研究院)
;
Shanghai Qi Zhi Institute(上海启智研究院)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Pennsylvania(宾夕法尼亚大学)
机构
*
Thomas Lord Department of Computer Science, University of Southern California(汤姆·劳德计算机科学系,南加州大学)
;
Meta AI
;
Department of Computer Science & Engineering, University of California San Diego(计算机科学与工程系,加州大学圣地亚哥分校)
Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning
Rewind-IL:在线故障检测与状态重置用于模仿学习
Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
机构
*
College of Connected Computing, Vanderbilt University(范德比尔特大学连接计算学院)
;
School of Computer Science, University of Waterloo(滑铁卢大学计算机科学学院)
;
School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
;
Australian Centre for Robotics, The University of Sydney(悉尼大学机器人研究中心)
机构
*
Department of Applied Mathematics, Ferdowsi University of Mashhad(菲尔多西大学应用数学系)
;
Department of Electronics, Information and Bioengineering (DEIB), Politecnico di Milano(米兰理工大学电子、信息与生物工程系)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
通过Stackelberg近端策略优化实现高效的形态-控制协同设计
Yanning Dai, Yuhui Wang, Dylan R. Ashley, Jürgen Schmidhuber
机构
*
Center of Excellence for Generative AI, King Abdullah University of Science and Technology (KAUST)(生成人工智能卓越中心,卡奥斯特大学)
;
Dalle Molle Institute for Artificial Intelligence Research (IDSIA)(人工智能研究达勒莫利研究所)
;
Università della Svizzera italiana (USI)(瑞士意大利大学)
;
Scuola universitaria professionale della Svizzera italiana (SUPSI)(瑞士意大利专业大学)
Commentspresented at the Fourteenth International Conference on Learning Representations; 11 pages in main text + 3 pages of references + 23 pages of appendices, 5 figures in main text + 11 figures in appendices, 16 tables in appendices; accompanying website available at https://yanningdai.github.io/stackelberg-ppo-co-design/ ; source code available at https://github.com/YanningDai/StackelbergPPO
RACAS: Controlling Diverse Robots With a Single Agentic System
RACAS:通过单一代理系统控制多样化机器人
Dylan R. Ashley, Jan Przepióra, Yimeng Chen, Ali Abualsaud, Nurzhan Yesmagambet, Shinkyu Park, Eric Feron, Jürgen Schmidhuber
机构
*
Center of Excellence in Generative AI, King Abdullah University of Science and Technology (KAUST), Saudi Arabia(沙特王国科学与技术大学生成人工智能卓越中心)
;
Dalle Molle Institute for Artificial Intelligence Research (IDSIA), Switzerland(人工智能研究达勒莫利 institute)
;
Università della Svizzera italiana (USI), Switzerland(瑞士意大利大学)
;
Scuola universitaria professionale della Svizzera italiana (SUPSI), Switzerland(瑞士意大利专业大学)
;
Robotics, Intelligent Systems, and Control Lab, King Abdullah University of Science and Technology (KAUST), Saudi Arabia(机器人、智能系统与控制实验室,沙特王国科学与技术大学(KAUST))
;
Department of Process Control, AGH University of Krakow, Poland(波兰克拉科夫AGH大学过程控制系)
Comments7 pages in main text + 1 page of appendices + 1 page of references, 5 figures in main text + 1 figure in appendices, 2 tables in main text; source code available at https://github.com/janprz11/robot-agnostic-control
机构
*
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室)
Integrating LTL Constraints into PPO for Safe Reinforcement Learning
将LTL约束整合到PPO中以实现安全强化学习
Maifang Zhang, Hang Yu, Qian Zuo, Cheng Wang, Vaishak Belle, Fengxiang He
机构
*
School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)
;
School of Computer Science, Faculty of Engineering, University of Sydney(计算机科学学院,工程学院,悉尼大学)
;
School of Engineering and Physical Sciences, Heriot-Watt University(工程与物理科学学院,赫瑞-瓦德大学)
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
;
National University of Singapore(新加坡国立大学)
机构
*
Dept. of Computer Science University of Minnesota, Twin Cities(计算机科学系明尼苏达大学双城分校)
;
Depts. of Computer Science and Psychology University of Minnesota, Twin Cities(计算机科学与心理学系明尼苏达大学双城分校)