Non-myopic Generation of Language Models for Reasoning and Planning
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
Comments Working in progress
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
Comments Work in progress
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG
Comments DPhil Thesis - Engineering Science, University of Oxford. Original copy available at https://ora.ox.ac.uk/objects/uuid:19489c19-dc5a-464a-831d-bbf887687c41
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI、cs.LG
Comments Presented at IEEE Conference on Robotics and Automation (ICRA) 2024. Website: https://sites.google.com/view/rdmemory
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
Comments 13 pages,6 figures,38 references
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2023 Track on Datasets and Benchmarks
评估VLMs在机器人运动中的空间推理能力:迈向具有运动偏好的机器人规划的一步
机构 * King’s College London(伦敦国王学院) ; University College London(伦敦大学学院)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文评估了四种先进VLMs在机器人运动中的空间推理能力,探讨了运动偏好(如物体接近性和路径风格)的处理,并分析了准确率与计算成本的权衡。
Comments Accepted to the First Workshop on Efficient Spatial Reasoning at ICLR 2026
Self-CriTeach: LLM 自我教学与自我批评用于通过自动领域生成提升机器人规划
机构 * Huawei Noah's Ark Lab(华为诺亚实验室) ; University of Toronto(多伦多大学) ; University of British Columbia(不列颠哥伦比亚大学) ; McGill University(麦吉尔大学)
专题命中 规划推理 :planning(title,abstract);CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)
AI总结 本文提出Self-CriTeach框架,通过LLM自动生成符号规划领域,用于自我教学生成规划问题-计划对及自我批评生成结构化奖励信号,提升机器人规划性能与泛化能力。
Comments International Conference on Machine Learning (ICML) 2026
ThinkingVLA:用于机器人操作的交叉视觉与语言推理
机构 * Fudan University(复旦大学)
专题命中 规划推理 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);planning(abstract)
AI总结 提出ThinkingVLA,通过统一的多Transformer架构实现前向与逆向推理交织,显著提升长时域操作任务性能。
GeoSolver: 利用细粒度过程监督扩展遥感中的测试时推理
机构 * College of Computer Science and Technology(计算机科学与技术学院) ; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education Jilin University(教育部符号计算与知识工程重点实验室)
专题命中 规划推理 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);verifier(abstract)
AI总结 提出GeoSolver框架,通过构建大规模过程监督数据集Geo-PRM-2M和训练过程奖励模型GeoPRM,结合过程感知树GRPO强化学习算法,实现遥感中可验证的逐步推理,在多个基准上达到最优性能并支持测试时扩展。
Comments Code: https://github.com/yourname/GeoSolver
隐秘动机:在连续思维模型中检测不一致的推理
机构 * Stanford University(斯坦福大学)
专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)
AI总结 本文研究了连续思维模型中如何检测不一致的推理,提出MoralChain基准测试,并通过双触发方法训练模型,揭示了潜在空间中不一致推理的存在及早期规划阶段的安全监控重要性。
Comments 15 pages with 2 figures
Journal ref International Conference on Learning Representations (ICLR) Latent & Implicit Thinking (LIT) Workshop 2026
思维分支:解释大语言模型推理需要重采样
机构 * MATS
专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);planning(abstract)
AI总结 本文通过重采样研究大语言模型推理过程中的因果影响,探讨了单个推理链的局限性,并提出了基于重采样的方法来分析模型决策和推理步骤的影响。
Comments Uzay Macar and Paul C. Bogdan contributed equally to this work, and their listed order was determined by coinflip
深度上限:大语言模型在发现潜在规划中的局限性
机构 * University of Cambridge(剑桥大学) ; Imperial College London(帝国理工学院) ; MIT(麻省理工学院)
专题命中 规划推理 :planning(title,abstract);reasoning(abstract);chain-of-thought(abstract);CoT(abstract)
AI总结 研究探讨了大语言模型在无监督情况下发现多步规划策略的极限,发现模型在单次前向传递中能执行最多七步潜在规划,揭示了训练与测试时能力的差异。
Comments 10 pages, 3 figures, 1 table (30 pages, 9 figures, 10 tables including references and appendices)
DraCo: 文本到图像预览与稀有概念生成的草稿作为CoT
机构 * CUHK MMLab(香港中文大学 MMLab) ; CUHK IMIXR(香港中文大学 IMIXR) ; Sun Yat-Sen University(中山大学) ; SCUT(华南理工大学) ; CUHK (Shenzhen)(香港中文大学(深圳))
专题命中 规划推理 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);planning(abstract)
AI总结 DraCo通过结合文本和视觉内容的交错推理,提升文本到图像生成的精度和稀有概念生成能力。
Comments Project Page: https://github.com/CaraJ7/DraCo
大语言模型可以在过程监督下学习和泛化隐写链式思维
机构 * Mentorship for Alignment Research Students (MARS)(对齐研究 mentorship 项目) ; University College London(伦敦大学学院) ; Queen Mary University of London(伦敦女王学院) ; ML Alignment & Theory Scholars (MATS)(对齐与理论学者) ; Meridian Impact, Cambridge(剑桥 Meridian Impact) ; Universidad Carlos III de Madrid(马德里卡洛斯三世大学) ; Geodesic Research and University of Cambridge(Geodesic Research 和剑桥大学)
专题命中 规划推理 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);planning(abstract)
AI总结 大语言模型在过程监督下能够学习并泛化隐写链式思维,通过替换特定字符串实现推理编码,提升监控可靠性。
Comments 10 pages main text, 3 figures main text, 17 pages supplementary material, 1 figure supplementary material, accepted at NeurIPS 2025
机构 * CUHK MMLab(CUHK多媒体实验室) ; CUHK MiuLar Lab(CUHK MiuLar实验室) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 规划推理 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);planning(abstract)
Comments Project Page: https://github.com/CaraJ7/T2I-R1
SBCO:面向规划智能体的自监督、验证器驱动的工具链优化器
专题命中 规划推理 :planning(title,abstract);verifier(title,abstract);分类 cs.AI
AI总结 本研究针对规划任务提出SBCO,一种自监督的验证器驱动工具链优化器,其性能优于或相当的定制基线,且计算成本仅为其1/4至1/5.5。
超越排行榜:大语言模型智能体中工具使用、规划和推理失败的综合分析
机构 * University of Oxford(牛津大学)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文综合多篇论文形成大语言模型智能体局限性交叉分类法,识别出六个失败集群,通过迭代分组得出分类法,发现失败与任务长度非线性相关,单个子任务性能不能保证端到端成功,额外脚手架效果不佳,同时在部分任务上有进展。
Comments 16 pages, 3 tables, 1 figure
GRASP:用于高保真相关工作生成的图推理辅助综述规划
机构 * Department of Computer Science(计算机科学系)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL
AI总结 研究如何写文献综述,核心方法是结合大语言模型规划与图算法,主要贡献是提出GRASP框架,其两层图结构及拓扑感知剪枝能生成与人工撰写匹配的相关工作部分。
Comments 23 pages, 3 figures. Published in Findings of the Association for Computational Linguistics: ACL 2026
Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 36427-36449, San Diego, California, United States. Association for Computational Linguistics
理性逆推理:通过规划推断意图进行少样本模仿
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 研究人类能从少量示范学习新操作任务,而现代机器人模仿学习需大量示范且适应性差的差距。提出理性逆推理,将少样本模仿视为对潜在解释程序的推理,通过视觉语言模型和分层规划器迭代优化候选集,在少样本设置中表现优于基线。
LLM引导的规划:多模态核监管文档的多跳推理
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 提出LLM引导的规划方法,通过动态知识图谱状态和文档树工具,在多跳推理任务中实现81.5%准确率,显著优于无状态规划方法。
Comments Accepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond @ ICML 2026. 8 pages (main), 3 figures, 1 algorithm
PPA-Plan: 长上下文LLM推理中的前瞻性坑洞避免规划
机构 * Hanyang University(翰阳大学) ; University of California, Santa Barbara(加州大学圣芭芭拉分校)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.CL
AI总结 针对长上下文推理中计划生成不可靠的问题,PPA-Plan通过前瞻性策略预防逻辑错误,提升计划执行效果。
Comments Accepted to the Main Conference of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). 27 pages, 6 figures
模型空间推理作为反馈空间中的搜索用于规划领域生成
机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院) ; IBM Research(IBM研究院)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文探讨了通过反馈空间中的启发式搜索优化规划领域生成的质量,利用符号反馈机制提升生成效果。
Comments Accepted at ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling
WebUncertainty: 双层不确定性驱动的规划与推理用于自主网络代理
机构 * Hefei University of Technology(合肥工业大学) ; Academy of Cyber, CETC Group(中国电子科技集团 Cyber 学院)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文提出WebUncertainty框架,通过双层不确定性驱动的自适应规划和蒙特卡洛树搜索机制,提升自主网络代理在动态交互和长周期任务中的规划与推理能力。
多模态大语言模型中的忠实优先推理、规划与行动
机构 * Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong University of Science and Technology(香港科技大学) ; The Australian National University(澳大利亚国立大学) ; Fudan University(复旦大学) ; University of Sydney(悉尼大学)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文提出Faithful-First RPA框架,通过逐步监督提升多模态推理的忠实度,实验表明其在多个基准上提升了24%的感知忠实度,且不降低任务准确性。
Comments Accepted by ACL 2026 Findings
通过代理快慢规划弥合大模型推理与实时控制
机构 * Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong (Shenzhen)(深圳市大数据研究院,香港中文大学(深圳)) ; Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究院) ; State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau(澳门大学智慧城市物联网国家重点实验室)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI
AI总结 本文提出一种分层框架,通过感知-决策-轨迹-控制的分离,提升自动驾驶系统在扰动下的鲁棒性,减少侧向偏移和完成时间。
Comments 8 pages, 12figures
通过蒙特卡洛网络信息增益实现链式推理过程监督
机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) ; IBM Research(IBM研究院)
专题命中 规划推理 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL
AI总结 本文提出利用信息论自动生成推理步骤标签,提升大语言模型推理过程的监督效率与可靠性,适用于多种推理基准测试。
GNNVerifier:基于图的LLM任务规划验证器
机构 * Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 规划推理 :planning(title,abstract);verifier(title,abstract);分类 cs.LG
AI总结 本文提出基于图的验证器,通过构建任务计划的图结构,利用图神经网络评估计划的合理性,从而纠正LLM生成的计划中的错误。
Comments 17pages,12figures
NaviDriveVLM:解耦高层推理与运动规划用于自动驾驶
机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA(杰米·沃克’66机械工程系,德克萨斯A&M大学,学院站,TX 77843,美国) ; The Department of Engineering Technology and Industrial Distribution Texas A&M University, College Station, TX 77843, USA(工程技术与工业分配系,德克萨斯A&M大学,学院站,TX 77843,美国)
专题命中 规划推理 :reasoning(title,abstract);planning(title,abstract);分类 cs.LG
AI总结 NaviDriveVLM通过解耦高层推理与运动规划,提升自动驾驶中的推理能力与效率