Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
奖励驱动的大语言模型智能体工作流程:合成用于自主决策的部分可观测马尔可夫决策过程(POMDP)路由与自我修正
专题命中 规划决策 :agent(title,abstract);autonomous agent(abstract);planning(abstract);workflow(abstract)
AI总结 研究针对LLM智能体应用挑战,设计优化工作流程,合成多种AI范式,引入POMDP路由和自我修正奖励模型,整合多模态输入与强化学习原则,实验显示相比主流基线任务成功率等提升24.5%,为自主系统AI技术开发提供参考框架。
Comments 29 pages, 1 table, native TikZ/pgfplots diagrams. Code available at https://github.com/01Amez/RLAW_Implementation