arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1117 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1117 篇

2309.12941 2023-09-25 cs.SE cs.AI 79%

Trusta: Reasoning about Assurance Cases with Formal Methods and Large Language Models

Zezhong Chen, Yuxin Deng, Wenjie Du

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments 38 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.15118 2023-08-30 cs.CL 79%

Large Language Models on the Chessboard: A Study on ChatGPT's Formal Language Comprehension and Complex Reasoning Skills

Mu-Tien Kuo, Chih-Chung Hsueh, Richard Tzong-Han Tsai

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02324 2023-04-18 cs.AI cs.GT cs.MA 79%

Reasoning about Causality in Games

Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, Alessandro Abate, Michael Wooldridge

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments Published in Artificial Intelligence (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03070 2022-09-20 cs.AI 79%

An Argumentation-Based Legal Reasoning Approach for DL-Ontology

Zhe Yu, Yiwei Lu

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments 16 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09054 2021-12-17 cs.CL 79%

Pushing the Limits of Rule Reasoning in Transformers through Natural Language Satisfiability

Kyle Richardson, Ashish Sabharwal

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

Comments Accepted to AAAI-2022, AAAI preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12641 2021-10-14 cs.CL 79%

Negation in Cognitive Reasoning

Claudia Schon, Sophie Siebert, Frieder Stolzenburg

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

Comments 19 pages, 5 figures, 4 tables, extended version

Journal ref In Stefan Edelkamp, Ralf Moeller, and Elmar Rueckert, editors, KI 2021: Advances in Artificial Intelligence - 44th German Conference on AI, LNAI 12873, pages 217-232. Springer, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04165 2020-10-29 cs.LO cs.AI cs.PL 79%

Proof-Carrying Plans: a Resource Logic for AI Planning

Alasdair Hill, Ekaterina Komendantskaya, Ronald P. A. Petrick

专题命中 代码与定理证明 :planning(title,abstract);分类 cs.AI

Comments PPDP 2020, 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.12737 2020-05-27 cs.AI cs.LO 79%

Towards United Reasoning for Automatic Induction in Isabelle/HOL

Yutaka Nagashima

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments This is the pre-print of our short-paper accepted at the 34th Annual Conference of the Japanese Society for Artificial Intelligence, 2020 (https://www.ai-gakkai.or.jp/jsai2020/en)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.11666 2019-12-24 stat.ML cs.CV cs.LG 79%

Learning Dynamics of Attention: Human Prior for Interpretable Machine Reasoning

Wonjae Kim, Yoonho Lee

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.LG

Comments 20 pages, 18 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.04005 2016-10-14 cs.AI cs.NI 79%

Stream Reasoning-Based Control of Caching Strategies in CCN Routers

Harald Beck, Bruno Bierbaumer, Minh Dao-Tran, Thomas Eiter, Hermann Hellwagner, Konstantin Schekotihin

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.07398 2016-08-29 cs.LO cs.AI cs.SE 79%

Proceedings First Workshop on Causal Reasoning for Embedded and safety-critical Systems Technologies

Gregor Gössler, Oleg Sokolsky

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Journal ref EPTCS 224, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1407.5380 2014-07-22 cs.AI 79%

Representing and Reasoning about Game Strategies

Dongmo Zhang, Michael Thielsher

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1404.5643 2014-04-24 cs.AI cs.MA 79%

A Formal Analysis of Required Cooperation in Multi-agent Planning

Yu Zhang, Subbarao Kambhampati

专题命中 代码与定理证明 :planning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1304.2361 2013-04-10 cs.AI 79%

Rational Nonmonotonic Reasoning

Carl Kadie

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments Appears in Proceedings of the Fourth Conference on Uncertainty in Artificial Intelligence (UAI1988)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18587 2026-06-01 cs.LG cs.AI cs.LO cs.PL 79%

Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs

编译以压缩:通过编译器输出提升形式定理证明器

Guchan Li, Rui Tian, Hongning Wang

机构 * Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)

专题命中 代码与定理证明 :reasoning(abstract);test-time compute(abstract);verifier(abstract);分类 cs.AI、cs.LG

AI总结 利用编译器将大量证明尝试压缩为结构化失败模式,提出一种学习-精炼框架,通过树搜索基于验证器反馈局部修正错误,在可比测试时预算下在PutnamBench上达到最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12603 2025-09-17 cs.CL cs.AI 79%

EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving

Mukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu

机构 * Tencent(腾讯) The University of Hong Kong(香港大学)

专题命中 代码与定理证明 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14026 2025-03-24 cs.CL cs.AI 79%

Uncovering Latent Chain of Thought Vectors in Language Models

Jason Zhang, Scott Viteri

专题命中 代码与定理证明 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments This work was presented at the Workshop on Neural Network Weights as a New Data Modality at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25996 2026-07-29 cs.SE 新提交 78%

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models

RepoReasoner:评估长上下文语言模型的仓库级代码推理能力

Yanlin Wang, Suiquan Wang, Yanli Wang, Bowen Zhang, Daya Guo, Jiachi Chen, Zibin Zheng

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 研究针对现有代码推理基准测试局限,引入RepoReasoner评估仓库级代码推理,通过多阶段管道构建基准,评估输出预测和调用链预测能力,发现当前LLMs在仓库级推理存在局限,为未来相关研究提供方向。

Comments Published in FSE'2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17981 2026-07-27 cs.SE 版本更新 78%

Planning to Hammer: Difficulty-Aware Decomposition for Automating Rocq Proofs

规划锤击:面向自动化Rocq证明的难度感知分解

Ning Zhang, Nongyu Di, Zenan Li, Yuan Yao, Xiaoxing Ma

专题命中 代码与定理证明 :planning(title,abstract)

AI总结 提出Quarry框架,通过LLM规划证明分解并利用难度模型排序子目标,结合CoqHammer自动证明,在Rocq基准测试中成功率提升7%-13%。

Comments 26 pages, 8 figures; submitted to OOPSLA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01223 2026-07-21 cs.AI cs.CL cs.LG cs.LO cs.SE 版本更新 78%

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

Theoria: 非正式推理状态上的重写-可接受性验证

Michael Saldivar, Ben Slivinski

机构 * Independent Researchers(独立研究者)

专题命中 代码与定理证明 :reasoning(title);分类 cs.CL、cs.AI、cs.LG

AI总结 提出Theoria验证架构,通过将候选解重写为带显式理由的序列化状态转换并验证变更完整性,在HLE-Verified Gold上以91.4%精确率认证105个问题,优于整体式LLM评判。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08458 2026-06-09 cs.RO 新提交 78%

Personalized and Robust Proactive Robot Assistance with Uncertainty-Guided LLM Reasoning

个性化且鲁棒的主动机器人辅助:基于不确定性引导的大语言模型推理

Alvaro Gonzalez, M. H. Hasan Shovo, Ali Ayub

机构 * Concordia University(康考迪亚大学)

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 提出GLOBE框架,结合n-gram马尔可夫模型与不确定性引导的大语言模型推理,在家庭环境中实现高效鲁棒的主动机器人辅助,并在HOMER-Noise数据集上验证了其性能与效率。

Comments Accepted to the 2026 IEEE 35th International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.02770 2026-06-04 cs.LO cs.SE cs.SY eess.SY 78%

Proceedings 2nd International Workshop on Causal Reasoning for Embedded and safety-critical Systems Technologies

第二届嵌入式和安全关键系统技术因果推理国际研讨会论文集

Alex Groce, Stefan Leue

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文探讨了复杂嵌入式和安全关键系统中因果推理的方法,旨在促进不同领域研究者之间的交流,以提升关键系统故障根本原因分析的科学性。

Journal ref EPTCS 259, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.04595 2026-06-02 econ.TH 78%

Measuring Information Burden: From Coalition-Based Reasoning to the Price System

衡量信息负担:从基于联盟的推理到价格体系

Shuige Liu

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文通过形式化框架衡量价格体系节省的信息量,证明在可转移效用博弈中,核心与一致同意等价所需的最小信息结构是每个联盟至少被其一个成员所知,且平均每代理信息负担随经济规模指数增长。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18384 2026-05-27 cs.RO cs.FL 78%

LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback

LAD-VF:LLM自动微分实现基于形式化方法反馈的无微调机器人规划

Yunhao Yang, Junyuan Hong, Gabriel Jacob Perin, Zhiwen Fan, Li Yin, Zhangyang Wang, Ufuk Topcu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of São Paulo(圣保罗大学) Texas A&M University(德克萨斯A&M大学) SylphAI

专题命中 代码与定理证明 :planning(title,abstract)

AI总结 提出LAD-VF框架,利用形式化验证反馈和LLM自动微分自动优化提示词,无需微调即可提升机器人规划任务中规范符合率,成功率从60%提升至90%以上。

Comments Presented at ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17914 2026-05-19 cs.PL 78%

Guiding LLM-based Loop Invariant Synthesis via Feedback on Local Reasoning Errors

通过反馈局部推理错误引导基于LLM的循环不变式合成

Tianchi Li, Zhenyu Yan, Junhao Liu, Peng Di, Xin Zhang

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文提出了一种框架,通过形式验证LLM的思考过程并检测局部推理错误,为基于LLM的循环不变式合成提供建设性反馈,从而提高合成效果。

Comments Accepted by ACM Transactions on Programming Languages and Systems (TOPLAS). DOI: 10.1145/3806652

Journal ref ACM Trans. Program. Lang. Syst. 48, 2, Article 8 (May 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12881 2026-04-15 cs.SE 78%

Evaluating LLMs Code Reasoning Under Real-World Context

评估大语言模型在现实世界上下文中的代码推理能力

Changshu Liu

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文提出R2Eval1基准测试,包含135个来自十个广泛使用的Python项目的代码推理问题,通过序列化复合和自定义类型,更真实地评估LLM的通用性。

Comments Accepted by ICES SRC (ACM Student Research Competition)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04500 2026-04-07 cs.CV 78%

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward

Saliency-R1: 通过显著图对齐奖励增强可解释性和忠实的视觉-语言推理

Shizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma, Qi Dou

机构 * The Chinese University of Hong Kong(香港中文大学) City University of Hong Kong(香港城市大学)

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 Saliency-R1通过显著图对齐奖励提升视觉-语言模型的推理可解释性和忠实性,利用显著图技术高效突出关键图像区域,结合GRPO优化策略增强模型对相关区域的关注,实验显示提升推理忠实性和任务性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25565 2026-03-27 cs.CV 78%

GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing

GeoHeight-Bench:迈向遥感中基于高度的多模态推理

Xuran Hu, Zhitong Xiong, Zhongcheng Hong, Yifang Ban, Xiaoxiang Zhu, Wufan Zhao

机构 * KTH Royal Institute of Technology(瑞典皇家理工学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Technical University of Munich(慕尼黑工业大学)

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文提出GeoHeight-Bench,通过构建高度感知的遥感理解评估框架,解决现有多模态模型在复杂遥感几何和灾害场景中忽略垂直维度的问题,提出数据生成管道和基准测试,并验证高度感知的重要性。

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23470 2026-03-25 cs.SE 78%

ConceptCoder: Improve Code Reasoning via Concept Learning

ConceptCoder: 通过概念学习提升代码推理

Md Mahbubur Rahman, Hengbo Tong, Wei Le

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 ConceptCoder通过学习代码概念提升代码推理能力,在漏洞检测任务中显著提高F1值,优于现有方法,并扩展至分支预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19858 2026-03-23 cs.RO cs.MA 78%

Beyond detection: cooperative multi-agent reasoning for rapid onboard EO crisis response

超越检测:用于快速在轨遥感危机响应的协作多智能体推理

Alejandro D. Mousist, Pedro Delgado de Robles Martín, Raquel Lladró Climent, Julian Cobos Aparicio

专题命中 代码与定理证明 :reasoning(title,abstract)

AI总结 本文提出了一种分层多智能体架构,用于在轨遥感处理,在资源和带宽受限条件下,通过协调专用AI智能体实现互补多模态观测的利用,减少计算开销并保持决策一致性。

Comments Accepted for presentation at the ESA's 4S Symposium 2026 Conference (see https://atpi.eventsair.com/4s-symposium-2026/)

详情

展开后加载摘要…

URL PDF HTML 收藏