Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
学习幅度而非方向:面向多轮多步骤大语言模型智能体的验证器约束信用分配
专题命中 测试时计算 :verifier(title,abstract);分类 cs.AI
AI总结 该研究提出CrEST分层信用分配框架,保留RL的验证器约束上限并结合自我教师的密集token级信号,在BFCL V3和WildToolBench上优于基线,可简化教师在策略优化中的作用。