Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning
Claw-R1:面向智能体强化学习的步骤级数据中间件系统
机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室)
专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);post-training(abstract);分类 cs.CL、cs.LG
AI总结 提出Claw-R1系统,通过网关服务器和数据池组件,将智能体交互步骤转化为结构化数据资产,支持实时检查、质量筛选和训练批次配置,解决智能体强化学习中数据生命周期管理问题。