LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models
专题命中 推理与问题求解 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 14 pages, 5 figures
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
专题命中 推理与问题求解 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 14 pages, 5 figures
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(title);prompting(abstract)
Comments Project Page: : https://vis-www.cs.umass.edu/3dllm/
专题命中 推理与问题求解 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
延迟验证破坏多智能体LLM信念:不稳定性阈值与最优校正器放置
机构 * Independent Researcher(独立研究者)
专题命中 推理与问题求解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
AI总结 研究多智能体LLM系统中延迟验证导致信念不稳定的问题,通过图论和谱分解推导出验证剂量的稳定性阈值,并提出贪心近似算法优化校正器放置。
Comments 20 pages, 5 figures, 1 table. Code and data: https://github.com/YehudaItkin/delayed-verification-llm
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Department of Computer Science, Tsinghua University(清华大学计算机系)
专题命中 推理与问题求解 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG
Comments Accepted to ICCV 2025. Codes and dataset are available at: https://github.com/duowuyms/OpenCATP-LLM
专题命中 推理与问题求解 :LLM(title,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI
Comments Project website: https://www.llm-reasoners.net/
面向智能制造中大型语言模型增强的多智能体强化学习(MARL)中心参考架构
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 本文针对智能制造自适应控制的耦合需求,提出以MARL为中心的三层参考架构,探讨LLM在MARL中的附着点,明确传统MARL与LLM组件的适配场景,为相关研究提供结构化框架。
图语言:知识图谱如何与大语言模型对话
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 研究针对知识图谱与大语言模型融合的挑战,提出GRALAN模型,通过关系令牌实现KG在LLM语义空间的适配,在问答任务中显著优于现有方法,为KG-LLM融合提供新范式。
Comments Accepted to ISWC 2025
结合大语言模型与符号推理,通过可解释知识库实现多机器人时序规划
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 本研究提出PLANTOR框架,结合LLM与符号推理,通过可解释知识库实现多机器人时序规划,在基准场景及多臂装配场景验证,可减少手动建模但需人工修正,主张混合工作流。
KoRe: 为大型语言模型设计的紧凑知识表示
机构 * University of Trento(特伦托大学)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 本文提出KoRe方法,通过将知识图谱的1跳子图编码为紧凑离散知识标记,并注入到LLM中,从而提升模型的知识推理能力并减少token使用量。
心电图大语言模型:基于心电图的心脏推理基础模型
专题命中 推理与问题求解 :LLM(title,summary_cn);foundation model(title);large language model(abstract);language model(abstract)
AI总结 研究针对一线分诊缺乏确定性成像及现有心电图人工智能系统局限的问题,提出ECG-LLM模型,用多模态到语言监督策略训练,能从心电图回答多样心血管问题,恢复测量值、预测表型,在相关任务中表现良好,为临床推理提供支持。
基于PDDLStream的任务与运动规划大语言模型的系统研究
机构 * Stony Brook University(石溪大学)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 研究大语言模型在机器人任务与运动规划中的应用,开发16种用LLMs替代关键TAMP组件的算法,通过实验发现基于LLM的规划器成功率低、规划时间长,提供几何细节会增加任务规划错误,多数情况下直接LLM变体表现更好。
Comments 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 8 pages, 6 figures
HRO:用于基于大语言模型的零样本目标导航的分层房间到物体框架
机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) ; Joint Research Laboratory for Embodied Intelligence, Xinjiang University(新疆大学具身智能联合研究实验室) ; Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(新疆大学丝绸之路多语言认知计算国际联合研究实验室)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 针对零样本目标导航问题,提出LLM驱动的分层房间到物体(HRO)框架,引导智能体粗到细探索导航至目标物体,实验表明该框架在Gibson和HM3D数据集上优于现有基于LLM的方法。
Comments Main paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2026 (IEEE SMC 2026)
AI能像城市规划师一样推理吗?基于专业判断的大语言模型基准测试
机构 * School of Architecture and Urban Planning, Shenzhen University(深圳大学建筑与城市规划学院) ; Shenzhen Key Laboratory of Urban Spatial Information and Intelligent Modeling(深圳市城市空间信息与智能建模重点实验室) ; Department of Urban Planning and Design, The University of Hong Kong(香港大学城市规划与设计系)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 提出UPBench框架,通过4×5知识支柱与认知水平矩阵评估25个LLM,发现模型在分析任务上优于事实回忆和综合判断,揭示了规划知识的制度依赖性。
Comments This paper has been withdrawn by the authors because the current version requires substantial revision and further validation before it can be considered a reliable representation of the work
AdaPlanBench: 在世界约束和用户约束下评估大语言模型智能体的自适应规划能力
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 针对现有基准未充分探索渐进揭示的双重约束下的自适应规划问题,提出动态交互基准AdaPlanBench,通过307个家务任务和可扩展的约束构建流程,评估LLM智能体在交互中根据反馈迭代调整计划的能力。
Comments COLM 2026
Adam的定律:大型语言模型中的文本频率定律
机构 * FaceMind Corporation(FaceMind公司) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);prompting(abstract)
AI总结 本文提出文本频率定律,通过优化文本频率提升LLM的提示和微调效果,并通过文本频率蒸馏和课程训练方法验证了其有效性。
Comments ACL 2026 Main Conference; The latest version
谜语谜题:测试大型语言模型和人类的灵活推理
机构 * Department of Psychology, Princeton University(普林斯顿大学心理学系) ; Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 通过谜语谜题范式,发现LLM在真实谜语上准确率高(84.9%),但在谜语谜题上表现差(50.7%),而人类相反;错误分析表明LLM更倾向于过度使用创造性推理。
超越价值基准:通过对称Q分类测量大型语言模型中的价值结构对齐
机构 * TJUNLP Lab, School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院TJUNLP实验室)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 提出基于Q方法的对称人机评估框架,通过140项道德陈述的强制分布排序,测量LLM价值结构对齐,发现跨家族异质性和局部错位。
Comments 32 pages, 8 figures, 16 tables; accepted to ACL 2026 Main Conference
人类级推理:大型语言模型在逻辑与抽象推理上的比较研究
机构 * Universidade Federal de Santa Catarina(联邦圣卡塔琳娜大学)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 本文通过八个定制推理问题比较了多个LLM在逻辑和抽象推理能力上的表现,揭示了模型在推理任务上的差异及不足。
Comments 12 pages
Journal ref Proceedings of the 2026 Computer on the Beach
用大型语言模型模拟学生的Java编程错误
机构 * University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);prompting(abstract)
AI总结 探索用LLM模拟学生编程错误,评估五种模型在多样性和对齐性上的表现,发现Claude Sonnet 4平衡最佳,且合成错误与真实错误难以区分。
利用大型语言模型发现专家级纳什均衡算法
机构 * CFCS, School of Computer Science, Peking University, Beijing, China(计算机科学系,北京大学,北京,中国) ; School of Computing and Data Science, The University of Hong Kong, Pokfulam, Hong Kong(计算与数据科学学院,香港大学,薄扶林,香港)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 提出LegoNE框架,将专家证明策略编码为符号语言,自动验证算法的最坏情况保证,结合推理型LLM重新发现并改进了多人博弈的近似纳什均衡算法。
Comments accepted by Nature Communications
赋予大语言模型双向逻辑以进行稳健的链修复
机构 * Department of Computer Science, University of Oxford, UK(英国牛津大学计算机系) ; FLock.io ; Institute of Logic and Computation, TU Wien, Austria(奥地利技术大学逻辑与计算研究所)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);SFT(abstract,abstract_cn)
AI总结 针对自回归链式推理中错误雪崩问题,提出Teleological Reasoning Infilling (TRI)框架,通过将错误推理段重构为填充中间任务并引入前缀-后缀-中间序列重排,结合符号验证器监督微调和直接偏好优化,实现仅修复受损段的高效链修复。
Comments 25 Pages
Journal ref In Proceedings of European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases 2026
图上的代码:通过大型语言模型在知识图谱上进行迭代式程序化推理
机构 * Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全重点实验室) ; Shandong University(山东大学) ; Shandong University-Weihai Research Institute of Industrial Technology(山东大学威海工业技术研究院)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 提出Code-on-Graph (CoG)框架,通过将知识图谱模式表示为Python类并生成可执行代码,解决现有LLM-KG集成中操作符不灵活和知识注入不可扩展的问题,在WebQSP、CWQ和GrailQA上提升高达10.5%。
基于偏好最大可满足性的大语言模型可靠推理
机构 * Artificial Intelligence Research Institute (IIIA) Consejo Superior de Investigaciones Científicas (CSIC)(人工智能研究所(IIIA)西班牙国家科学研究委员会(CSIC)) ; Department of Computer Science University of Oxford(计算机科学系牛津大学) ; Institut de Robòtica i Informàtica Industrial (IRI-CSIC-UPC)(机器人与信息工业研究所(IRI-CSIC-UPC))
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 提出一种混合推理方法,通过LLM生成代码将自然语言问题编码为偏好最大可满足性问题,由精确求解器求解并独立验证,显著提高可行性。
Comments 17 pages, 1 figure, 4 tables
使用大型语言模型生成鲁棒的优化模型组合
机构 * Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) ; Harvard University(哈佛大学)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 提出一种利用LLM作为随机生成器和推理评估器的统一框架,生成鲁棒的优化模型组合,并保证在生成器或评估器之一与人类偏好对齐时组合中包含高质量候选模型。
Comments Accepted at the ICML 2026 LM4Plan Workshop
基于大语言模型的累积推理
机构 * IIIS, Tsinghua University(清华大学人工智能研究院) ; Shanghai Qi Zhi Institute(上海启智研究院)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 本文提出了一种名为累积推理(CR)的框架,通过模拟人类的迭代和累积思维过程,增强大语言模型(LLM)的问题解决能力。CR通过分解任务、生成并验证中间推理步骤,构建动态有向无环图(DAG)来组成解决方案,从而在逻辑推理、24点游戏和数学问题等任务中取得了显著的性能提升。
Comments Published in Transactions on Machine Learning Research (TMLR). Project Page: https://github.com/iiis-ai/cumulative-reasoning
ClassInvGen:利用大语言模型进行类不变式合成
机构 * Stanford University(斯坦福大学) ; Microsoft Research(微软研究院) ; Purdue University(普渡大学)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结 本文提出ClassInvGen,利用大语言模型生成C++等主流语言的可执行类不变式及测试输入,优于纯LLM方法和Daikon等传统技术,提供基准测试和案例研究验证其有效性。
增强型大型语言模型用于网络搜索中动态内容过期预测
机构 * Baidu Inc.(百度公司)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 本文提出基于大型语言模型的查询感知动态内容过期预测框架,通过提取细粒度时间上下文和利用LLM推断查询特定的有效性范围,提升搜索结果的新鲜度和用户体验。
Comments Accepted at SIGIR 2026. Final version: https://doi.org/10.1145/3805712.3808457
LegalDrill: 以诊断驱动合成促进小型语言模型中的法律推理
机构 * Purdue University(普渡大学) ; Fidelity Investments(富达投资)
专题命中 推理与问题求解 :language model(title,abstract);small language model(title,abstract);SLM(abstract,abstract_cn);preference optimization(abstract)
AI总结 LegalDrill通过精细提示提取并迭代优化教师模型的推理轨迹,结合自我反思验证选择有效数据,提升小型语言模型的法律推理能力,无需稀缺专家标注。
Comments ACL 2026 Industry Track
为大语言模型中的战略推理进行前瞻性优化
机构 * Department of Computing, Hong Kong Polytechnic University(香港理工大学计算机系) ; Department of Language Science and Technology, Hong Kong Polytechnic University(香港理工大学语言科学与技术系) ; Apple(苹果公司) ; School of Design, Hong Kong Polytechnic University(香港理工大学设计学院) ; Research Institute for Quantum Technology, Hong Kong Polytechnic University(香港理工大学量子技术研究所) ; Department of Communication Science, Vrije Universiteit Amsterdam(阿姆斯特丹自由大学传播科学系)
专题命中 推理与问题求解 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 本文提出FoPO方法,通过整合对手建模原理提升大语言模型的战略推理能力,验证了其在不同规模和来源的LLM中的有效性及泛化能力。
Comments ACL 2026 Main Conference