Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
在知识图谱上探索:通过路径细化奖励建模激励大语言模型自主探索
机构 * Zhongguancun Laboratory, Beijing, China(中关村实验室,北京,中国) ; Department of Electronic Engineering, Tsinghua University, Beijing, China(电子工程系,清华大学,北京,中国) ; Institute for Network Sciences and Cyberspace, Tsinghua University, Beijing, China(网络科学与网络空间研究院,清华大学,北京,中国) ; Ant International, Ant Group, Hangzhou, Zhejiang, China(蚂蚁集团,杭州,浙江,中国)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL
AI总结 本文提出Explore-on-Graph框架,通过强化学习和路径细化奖励建模,使大语言模型在知识图谱上自主探索更广泛的推理空间,实现对分布外推理问题的高效泛化。
Comments Published as a conference paper at ICLR 2026