arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45051 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1129 篇

1903.03515 2019-04-23 cs.AI 57%

Learning $\textit{Ex Nihilo}$

Selmer Bringsjord, Naveen Sundar Govindarajulu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.01886 2019-02-07 cs.AI 57%

Situational Grounding within Multimodal Simulations

James Pustejovsky, Nikhil Krishnaswamy

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments AAAI-19 Workshop on Games and Simulations for Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.03846 2019-02-06 cs.CY cs.AI cs.HC cs.RO stat.ML 57%

"Dave...I can assure you...that it's going to be all right..." -- A definition, case for, and survey of algorithmic assurances in human-autonomy trust relationships

Brett W Israelsen, Nisar R Ahmed

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments final version of accepted manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.09125 2019-01-29 cs.AI cs.LO 57%

The informal semantics of Answer Set Programming: A Tarskian perspective

Marc Denecker, Yuliya Lierler, Miroslaw truszczynski, Joost Vennekens

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.08762 2017-07-28 cs.AI cs.LO 57%

Argument-based Belief in Topological Structures

Chenwei Shi, Sonja Smets, Fernando R. Velázquez-Quesada

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments In Proceedings TARK 2017, arXiv:1707.08250

Journal ref EPTCS 251, 2017, pp. 489-503

详情

展开后加载摘要…

URL PDF HTML 收藏
1402.5043 2014-02-21 cs.AI 57%

A logical model of Theory of Mind for virtual agents in the context of job interview simulation

Marwen Belkaid, Nicolas Sabouret

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1310.6429 2013-10-28 cs.AI cs.LO 57%

Knowledge-Based Programs as Plans: Succinctness and the Complexity of Plan Existence

Jerome Lang, Bruno Zanuttini

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Comments 10 pages, Contributed talk at TARK 2013 (arXiv:1310.6382) http://www.tark.org

详情

展开后加载摘要…

URL PDF HTML 收藏
1307.2191 2013-07-09 cs.HC cs.AI 57%

A Knowledge-based Treatment of Human-Automation Systems

Yoram Moses, Marcia K. Shamo

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 39 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.3876 2013-01-18 cs.AI 57%

Probabilistic Models for Agents' Beliefs and Decisions

Brian Milch, Daphne Koller

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Appears in Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence (UAI2000)

详情

展开后加载摘要…

URL PDF HTML 收藏
1111.0041 2011-11-02 cs.AI cs.MA cs.PL 57%

On the Formal Semantics of Speech-Act Based Communication in an Agent-Oriented Programming Language

R. H. Bordini, A. F. Moreira, R. Vieira, M. Wooldridge

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 29, pages 221-267, 2007

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.1314 2011-09-08 cs.AI 57%

Measuring Intelligence through Games

Tom Schaul, Julian Togelius, Jürgen Schmidhuber

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1106.4867 2011-06-27 cs.AI 57%

Compiling Causal Theories to Successor State Axioms and STRIPS-Like Systems

F. Lin

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 19, pages 279-314, 2003

详情

展开后加载摘要…

URL PDF HTML 收藏
1012.1648 2010-12-09 cs.AI cs.CE 57%

Analysis Of Cancer Omics Data In A Semantic Web Framework

Matt Holford, James McCusker, Kei Cheung, Michael Krauthammer

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments in Adrian Paschke, Albert Burger, Andrea Splendiani, M. Scott Marshall, Paolo Romano: Proceedings of the 3rd International Workshop on Semantic Web Applications and Tools for the Life Sciences, Berlin,Germany, December 8-10, 2010

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16978 2026-04-07 cs.MA 56%

Lark: Biologically Inspired Neuroevolution for Multi-Stakeholder LLM Agents

Lark:生物启发的多利益相关者大语言模型代理神经进化

Rikhil Tanugula, Dheeraj Chintapalli, Sunkalp Chandra

专题命中 代码与定理证明 :reasoning(abstract,comments)

AI总结 Lark通过结合大语言模型推理与进化型多智能体系统,解决冗余与利益相关者权衡问题,采用四机制提升策略生成效率与透明度,实验显示其在30轮评估中表现优异且成本可控。

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20673 2026-04-03 cs.CL cs.AI 54%

PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs

PAVE:基于前提的验证与编辑用于检索增强的大语言模型

Tianyi Huang, Caden Yang, Emily Yin, Eric Wang, Michael Zhang

机构 * Ryquo App-In Club

专题命中 代码与定理证明 :分类 cs.CL、cs.AI;reasoning(comments);logical reasoning(comments)

AI总结 PAVE通过在推理阶段验证和编辑检索到的证据,提高检索增强大语言模型的回答一致性。在两个证据基础问答任务中,PAVE在跨度基础基准上提升了32.7个准确率点。

Comments Accepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19626 2026-08-21 cs.SE 新提交 50%

Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem

在神谕问题下对LLM测试生成中反馈驱动演化的审计与分解

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

专题命中 代码与定理证明 :verifier(abstract)

AI总结 该研究针对LLM测试生成中反馈驱动演化的神谕问题,通过多任务实验发现单一神谕会膨胀演化增益,提出审计-安慰剂协议以分离验证器人工制品等,为相关评估提供改进方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19240 2026-08-21 math.CV 新提交 50%

A Nonmonotone Real-Rootedness Set for Symmetric Imaginary Shifts

Vasily Stodolsky

专题命中 代码与定理证明 :verifier(abstract)

Comments 7 pages, 1 table. AI Research Artifact with material AI use and author accountability disclosed in the paper. Corresponds to Zenodo v0.1.3, DOI: 10.5281/zenodo.21922629

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18118 2026-08-20 math.OC cs.SY eess.SY 新提交 50%

Formal Safety Verification for Nonlinear Systems with Generative Barrier Certificate

基于生成式障碍证书的非线性系统形式化安全验证

Mengxin Ren, Hanrui Zhao

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 该研究针对非线性系统安全验证中障碍证书推导计算成本高的问题,提出基于大语言模型的生成式框架,将BMI问题转化为LMI测试,实现了远超传统方法的速度与性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21339 2026-08-17 cs.SE cs.PL 版本更新 50%

KBSpec: LLM-driven Formal Specification Generation with Evolving Domain Knowledge Base

KBSpec:基于演化领域知识库的LLM驱动形式化规约生成

Wenhan Wang, Zeyu Sun

专题命中 代码与定理证明 :verifier(abstract)

AI总结 提出KBSpec方法,利用外部官方文档和内部验证器反馈的双源知识增强LLM,通过自演化知识库持续更新成功轨迹,无需调参或标注数据,在JML规约生成上验证通过率提升10-25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13077 2026-08-14 cs.SE 新提交 50%

How Powerful are LLMs in Generating Formal Program Specifications?

大型语言模型在生成形式化程序规约方面的能力有多强?

Fanpeng Yang, Xing Li, Shuling Wang, Jie An, Zeyu Sun, Shenghua Feng, Wenhan Wang, Weiyi Wang, Naijun Zhan, Fanjiang Xu

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 该研究引入基于Rocq的Coins评估框架,在HumanEval数据集上开展大规模研究,发现LLMs生成形式化程序规约仍具挑战,准确的规约评估是理解其能力的核心。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11394 2026-08-13 cs.SE 新提交 50%

GraphAlignCoder: Aligning Program and Proof Graphs for Code Generation

GraphAlignCoder:对齐程序与证明图的代码生成框架

Yueke Zhang, Zihan Fang, Kevin Leach, Yu Huang

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 GraphAlignCoder是将显式正确性结构迁移至代码生成的训练框架,通过对齐程序与证明图提升性能,在多个基准测试中优于现有方法,验证图注入与验证到代码的整合是关键。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09312 2026-08-11 cs.NI 新提交 50%

Automated Synthesis of Deterministic Cross-Domain Interfaces

确定性跨域接口的自动合成

Konstantinos Christodoulopoulos, Antonis Selentis-Boulntadakis

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 该研究提出框架,利用LLM实现智能体与处理程序模块,自动合成跨域确定性网络的静态、动态契约,经实验验证其动态契约性能优于静态配置,且分工模式可保障模型可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17798 2026-08-11 cs.SE cs.CR 版本更新 50%

CognixShield: PoV-Guided Vulnerable API Usage Detection in Large Codebases via LLMs

CognixShield:基于LLM的、面向大型代码库的PoV引导式漏洞API使用检测

Quanzhi Fu, Wang Lingxiang, Wenjia Song, Gelei Deng, Yi Liu, Dan Williams, Ying Zhang

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 CognixShield是一款LLM驱动的框架,通过保留语义的AST碎片化、感知漏洞的多智能体RAG、PoV引导的语义推理三个核心组件,在57个Java应用上实现了优于现有工具的漏洞API使用检测性能。

Comments accepted in ESEM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06166 2026-08-07 cs.CY 新提交 50%

What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

开箱即用的大语言模型(LLM)在法律领域能(不能)做什么?针对意大利律师、法官和公证人考试的图灵测试

Germana Bertoli, Ilaria Amelia Caggiano, Francesca Lagioia, Riccardo Rovatti, Giovanni Sartor, Emiliano Troisi

专题命中 代码与定理证明 :planning(abstract)

AI总结 本研究通过意大利律师、法官、公证人考试的盲法图灵测试,评估开箱即用的主流LLM的法律能力,发现其在部分法律任务表现接近人类但公证人考试全部失败,明确了LLM法律能力的范围与边界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26368 2026-08-06 cs.NI 版本更新 50%

Introducing Large Language Models into the Design Flow of Time-Sensitive Networking

将大语言模型引入时间敏感网络的设计流程

Rubi Debnath, Luxi Zhao, Mohammadreza Barzegaran, Paul Pop, Sebastian Steinhorst

专题命中 代码与定理证明 :planning(abstract)

AI总结 针对配置优化TSN网络的挑战,通过跨模型案例研究评估现有大语言模型能力,提出LLM辅助编排框架,介绍构建模块与管道,分析实际部署机会与局限,为评估LLM辅助TSN编排可行性提供路线图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01938 2026-08-04 cs.CR 新提交 50%

D-MUTRA: DLT-based MUTual Remote Attestation for Multi-Agent Systems

D-MUTRA:面向多智能体系统的基于分布式账本(DLT)的相互远程证明

Adam Zahir, Vincent Lefebvre, Mark Angoustures, Milan Groshev, Carlos J. Bernardos

专题命中 代码与定理证明 :verifier(abstract)

AI总结 针对多智能体系统传统远程证明的局限,本文提出基于区块链的D-MUTRA框架,实现智能体间的持续相互证明,可检测恶意软件修改,且扩展时开销可忽略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27259 2026-07-31 cs.LO cs.AR 新提交 50%

CircuitProver: Agentic Lean 4 Theorem Proving with Reusable Circuit Proof Library for Hardware Verification

CircuitProver:基于可复用电路证明库的智能体Lean 4定理证明框架,用于硬件验证

Ziyi Yang, Wenji Fang, Chen Chen, Zhiyao Xie, Hongce Zhang

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 CircuitProver是基于Lean 4的智能体硬件验证框架,可自动转换硬件设计为Lean模型,通过积累可复用证明知识完成63项基准证明,比普通智能体更高效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17477 2026-07-28 math.GR math.CO 版本更新 50%

On Some Problems from the Kourovka Notebook

关于《库罗夫卡笔记本》中的一些问题

Wouter van Doorn, Elias Judin, Pietro Monticone, Daniel Morrison

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 本文解决《库罗夫卡笔记本》中八个群论问题,包括构建特定群、证明元素乘积取值情况、说明群阶与统计量不能确定单性、构造算子,还涉及确定生成群、证明幂图性质、探讨子群格情况及反驳秩不等式,且由形式推理主体在Lean中完成。

Comments 21 pages. Lean 4 formalisation: https://github.com/pitmonticone/Kourovka. v2: Expands the appendix with an account on problem selection and adds the MathOverflow provenance of Problem 18.50

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02124 2026-07-23 cs.NI cs.ET 版本更新 50%

FlexNGIA 2.0: Redesigning the Internet with Agentic AI -- Protocols, Services, and Traffic Engineering Designed, Deployed, and Managed by AI

FlexNGIA 2.0:利用智能AI重新设计互联网——由AI设计、部署和管理的协议、服务和流量工程

Mohamed Faten Zhani, Younes Korbi, Yamen Mkadem

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 面对沉浸式通信需求等带来的契机,FlexNGIA 2.0利用基于大语言模型的AI智能体自主编排、配置和演进网络,可动态调整多种网络方案。通过初步实验证明其能力,为新型智能AI驱动网络奠基,也指出了相关关键研究挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18727 2026-07-22 cs.PL cs.AR 新提交 50%

Formal Verification of an Out-of-Order Multiprocessor against an In-Order Weak-Memory ISA

针对顺序弱内存ISA的乱序多处理器形式验证

Janggun Lee, Jeehoon Kang

专题命中 代码与定理证明 :reasoning(abstract)

AI总结 研究针对顺序弱内存ISA的乱序多处理器验证问题,核心方法是设计核心规范并分两步证明,主要贡献是首次实现此类形式验证,借助核心规范简化证明,且利用LLM代理自动编写证明。

详情

展开后加载摘要…

URL PDF HTML 收藏