arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45006 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1127 篇

2503.04740 2025-03-10 cs.CY cs.AI cs.LG 81%

PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment

Anthony Diamond

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments 104 pages, 5 figures. Preprint on AI alignment presenting PRISM: a multi-perspective framework that organizes moral concerns into seven basis worldviews and uses Pareto-inspired synthesis to reconcile conflicting human values and specification gaming. Grounded in cognitive science and moral psychology

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18548 2025-02-26 cs.LO cs.AI cs.CC cs.LG 81%

Transformer Encoder Satisfiability: Complexity and Impact on Formal Reasoning

Marco Sälzer, Eric Alsmann, Martin Lange

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16075 2024-12-23 cs.AI cs.LG cs.LO 81%

Formal Mathematical Reasoning: A New Frontier in AI

Kaiyu Yang, Gabriel Poesia, Jingxuan He, Wenda Li, Kristin Lauter, Swarat Chaudhuri, Dawn Song

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15296 2024-12-23 cs.CL cs.LG 81%

Confidence in the Reasoning of Large Language Models

Yudi Pawitan, Chris Holmes

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03778 2024-07-08 cs.AI cs.CL cs.HC 81%

From Data to Commonsense Reasoning: The Use of Large Language Models for Explainable AI

Stefanie Krause, Frieder Stolzenburg

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13036 2024-05-24 cs.CL cs.AI 81%

Can formal argumentative reasoning enhance LLMs performances?

Federico Castagna, Isabel Sassoon, Simon Parsons

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01108 2024-02-05 cs.CL cs.LG 81%

Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions

Pouya Pezeshkpour, Eser Kandogan, Nikita Bhutani, Sajjadur Rahman, Tom Mitchell, Estevam Hruschka

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10663 2023-04-24 cs.CL cs.AI 81%

Meta Semantics: Towards better natural language understanding and reasoning

Xiaolin Hu

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages, 8 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10714 2022-05-25 cs.CL cs.AI 81%

Interpretable Proof Generation via Iterative Backward Reasoning

Hanhao Qu, Yu Cao, Jun Gao, Liang Ding, Ruifeng Xu

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments NAACL-HLT 2022 (Long), 14 pages (2 page references + 3 page appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08184 2019-08-23 cs.AI cs.LG 81%

Report on the First Knowledge Graph Reasoning Challenge 2018 -- Toward the eXplainable AI System

Takahiro Kawamura, Shusaku Egami, Koutarou Tamura, Yasunori Hokazono, Takanori Ugai, Yusuke Koyanagi, Fumihito Nishino, Seiji Okajima, Katsuhiko Murakami, Kunihiko Takamatsu, Aoi Sugiura, Shun Shiramatsu, Shawn Zhang, Kouji Kozaki

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08162 2019-01-25 cs.LG cs.AI stat.ML 81%

Causal Reasoning from Meta-reinforcement Learning

Ishita Dasgupta, Jane Wang, Silvia Chiappa, Jovana Mitrovic, Pedro Ortega, David Raposo, Edward Hughes, Peter Battaglia, Matthew Botvinick, Zeb Kurth-Nelson

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04452 2018-04-13 stat.ML cs.AI cs.LG 81%

Solving Bongard Problems with a Visual Language and Pragmatic Reasoning

Stefan Depeweg, Constantin A. Rothkopf, Frank Jäkel

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.4639 2013-11-20 cs.AI cs.LG cs.LO 81%

Post-Proceedings of the First International Workshop on Learning and Nonmonotonic Reasoning

Katsumi Inoue, Chiaki Sakama

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments 67 pages, 5 papers, 1 abstract, 1 cover

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13872 2026-05-15 cs.NE cs.AI 80%

S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning

S-AI-Recursive:一种生物启发式且时间稀疏的AI架构,用于迭代、反思和节能推理

Said Slaoui

机构 * Mohammed V University(穆莱·伊斯梅尔大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出S-AI-Recursive架构,通过激素闭环迭代而非单向传递实现推理,结合生物启发式方法和数学模型,验证了时间稀疏性原理。

Comments Preprint. 51 pages. No figures. S-AI-Recursive: A bio-inspired sparse AI architecture for iterative, introspective, and energy-efficient reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10563 2026-01-05 cs.AI 80%

NormCode: A Semi-Formal Language for Auditable AI Planning

NormCode: 一种用于可审计AI规划的半形式语言

Xin Guan, Yunshan Li, Zekun Wu, Ruibo Zhang

机构 * Center for Long-Term AI(长期人工智能中心) Shenzhen University(深圳大学) University College London(伦敦大学学院)

专题命中 代码与定理证明 :planning(title,comments);reasoning(abstract);分类 cs.AI

AI总结 NormCode是一种半形式语言,通过强制数据隔离和严格区分语义与语法操作,实现AI工作流的可审计性,确保透明性和可验证性。

Comments Archive name: NormCode: A Semi Formal Language for Context Isolated AI Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06101 2025-10-08 cs.CL 80%

The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models

Muyu He, Muhammad Ali Shafique, Anand Kumar, Tsach Mackey, Nazneen Rajani

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

Comments NeurIPS 2025 Workshop on Deep Learning for Code (DL4C), Project page: https://collinear.ai/valley-of-reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04412 2024-01-02 cs.AI cs.LO 80%

Human Conditional Reasoning in Answer Set Programming

Chiaki Sakama

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

Comments 34 pages. Shorter version: in Proceedings of the 21st International Workshop on Non-Monotonic Reasoning (NMR-2023)

Journal ref Theory and Practice of Logic Programming, vol.24(1), January 2024, pp. 157-192

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23916 2026-07-28 math.NA cs.NA 新提交 80%

Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems

递归治理:智能体人工智能系统中风险传播与漂移检测的图论框架

Sriram Nagaraj, Advaith Nila Narayanan

专题命中 代码与定理证明 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 针对金融机构模型风险管理问题,提出动态IaC治理循环,贡献包括引入校准DoA得分、构建DAG定义算法、开发轨迹监测协议及解决版本变化等实际复杂性问题,用于智能体人工智能系统风险传播与漂移检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25239 2026-05-13 cs.AI cs.CL cs.LG 80%

A Formal Comparison Between Chain of Thought and Latent Thought

链式思维与潜在思维的正式比较

Kevin Xu, Issei Sato

机构 * Department of Computer Science, The University of Tokyo, Japan(东京大学计算机科学系)

专题命中 代码与定理证明 :CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文对比了链式思维与潜在思维,发现潜在思维在并行计算上更高效,而链式思维能通过随机解码实现近似计数与采样。

Comments Camera-ready version for ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04910 2026-05-12 cs.SE 80%

Reducing the Costs of Proof Synthesis on Rust Systems by Scaling Up a Seed Training Set

通过扩大种子训练集来降低Rust系统证明合成的成本

Nongyu Di, Tianyu Chen, Shan Lu, Shuai Lu, Yeyun Gong, Peng Cheng, Jacob R. Lorch, Yuan Yao, Xiaoxing Ma

专题命中 代码与定理证明 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出VeruSyn数据合成管道,通过自合成和教程合成扩大Verus验证工具的数据集,生成690万Rust程序及其形式化规格证明,提升代码生成模型的性价比和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08957 2024-05-24 cs.AI cs.CL cs.FL cs.LG cs.PL 80%

MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data

Yinya Huang, Xiaohan Lin, Zhengying Liu, Qingxing Cao, Huajian Xin, Haiming Wang, Zhenguo Li, Linqi Song, Xiaodan Liang

专题命中 代码与定理证明 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref ICLR 2024 spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14178 2026-08-20 cs.AI cs.MA 版本更新 79%

ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System

ReasFlow:通过基于知识的多智能体系统助力应用数学中以推理为中心的科学发现

Yutong He, Daibo Li, Guohong Li, Jiahe Geng, Zhengyang Huang, Can Ren, Zekun Zhang, Yifan Liu, Shuchen Zhu, Hengrui Zhang, Boao Kong, Ming Sun, Shu Li, Chenyi Li, Jiang Hu, Kun Yuan, Zaiwen Wen, Pingwen Zhang

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 针对理论驱动科学发现探索不足的问题,ReasFlow引入以推理为中心的自主智能体系统,通过内部验证循环和知识检索机制减少专家干预,能统一多项科研任务,从最少提示生成高质量论文,在开源基线中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14566 2026-08-18 cs.AI 新提交 79%

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

立场:AI道德推理的评估仍遗漏了一半关键内容

Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh

机构 * University of Connecticut(康涅狄格大学) Carnegie Mellon University(卡内基梅隆大学) Hugging Face

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 该研究指出现有LLM道德推理评估聚焦道德价值问题而忽视规范问题,识别出三大缺口并提出含规范表征、标注数据集及分层评估的研究议程,以推动规范推理的系统研究。

Comments 8 pages, 1 figure. Accepted for archival publication at the ACL 2026 Workshop on Evaluating Evaluations (EvalEval)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12876 2026-08-14 cs.CV cs.AI 新提交 79%

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

SPARED:基于推理的AI生成图像检测方法,采用对抗编辑数据

Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 本研究提出对抗强化学习框架SPARED,通过扩散图像编辑器与推理型MLLM的交替博弈,训练出能抗捷径、泛化能力强的AI生成图像检测器,在三个外部基准上性能单调提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09277 2026-08-11 cs.AI cs.PL 新提交 79%

P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation

P³:用于验证代码生成的程序与证明联合规划

Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li, Kaiyu Yang

机构 * Apodex Princeton University(普林斯顿大学) Caltech(加州理工学院) University of Toronto(多伦多大学)

专题命中 代码与定理证明 :planning(title,abstract);分类 cs.AI

AI总结 P³是一种用于验证代码生成的程序与证明联合规划的LLM智能体工作流,在三个基准上的求解率优于基线,还降低了API成本与运行时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07593 2026-08-11 cs.HC cs.AI cs.IR 新提交 79%

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

感知天气与位置的智能体餐饮推荐:利用大语言模型世界知识开展区域敏感语境推理

Kadharmoideen Fadurudeen

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 该研究提出感知天气与位置的智能体餐饮推荐系统,利用LLM的世界知识生成区域敏感的适配天气的餐饮推荐,为智能体推荐提供了可扩展的架构模式。

Comments 5 pages. An agentic LLM system that reasons over combined location and weather context for region-sensitive dining recommendation. Working prototype implemented and briefly deployed end-to-end

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13077 2026-08-04 cs.MA cs.AI 版本更新 79%

Counterfactual Reasoning for Causal Responsibility Attribution in Probabilistic Multi-Agent Systems

反事实推理在概率多智能体系统中的因果责任归因

Chunyan Mu, Muhammad Najib

机构 * University of Aberdeen(阿伯丁大学) Heriot-Watt University(赫瑞瓦特大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出基于反事实推理的因果责任归因方法,利用Shapley值确保公平性和一致性,构建了支持责任感知多智能体系统验证与策略推理的框架,并通过纳什均衡计算稳定策略配置。

Comments Accepted at IJCAI 2026. This is the full version containing all proofs

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18443 2026-07-22 cs.CL 新提交 79%

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

具有灵活生成意义和表达替代方案的语用推理计算模型

Polina Tsvilodub, Fausto Carcassi, Michael Franke

机构 * University of Tübingen(图宾根大学) University of Amsterdam(阿姆斯特丹大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

AI总结 研究提出SAGE框架,结合认知模型与语言模型,将语用过程分解为提议者、评估者和选择器模块。通过三个案例研究评估,该模型取得高精度且常超基线,不过组件分析有不对称性,还探讨了神经符号模型在语用语言使用中的前景与局限。

Comments 27 pages main text, 9 figures; 25 pages supplementary materials

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03245 2026-07-22 cs.AR cs.AI cs.SE 版本更新 79%

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification

FVRuleLearner:基于操作级推理树(OP-Tree)的规则学习用于形式验证

Lily Jiaxin Wan, Chia-Tung Ho, Yunsheng Bai, Cunxi Yu, Ghaith Bany Hamad, Deming Chen, Haoxing Ren

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NVIDIA(英伟达) University of Maryland, College Park(马里兰大学帕克分校)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出FVRuleLearner,通过操作级推理树模型,提升形式验证中SVA操作符选择的准确性和效率,显著提高语法和功能正确性。

Comments Accepted to IEEE VTS'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11258 2026-07-14 cs.CL 新提交 79%

TreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMs

TreeThink:用于大语言模型数学推理的模块化树搜索库

Burak S. Akbudak, Zeynel A. Uluşan, Can S. Erer, Gözde Gül Şahin

机构 * Bogazici University(博阿齐奇大学) Codeway Studios(Codeway工作室) Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡大学) Koç University(科克大学) KUIS AI Lab(KUIS人工智能实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL

AI总结 介绍开源Python库TreeThink,用于神经定理证明的模块化全异步树搜索,集成多种方法技术,支持多种语言,连接REPL服务器,经评估在miniF2F和MATH500上有跨语言证明搜索等优势及异步加速。

详情

展开后加载摘要…

URL PDF HTML 收藏