arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45051 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1129 篇

2507.20199 2025-08-14 cs.AI 57%

StepFun-Prover Preview: Let's Think and Verify Step by Step

Shijie Shang, Ruosi Wan, Yue Peng, Yutong Wu, Xiong-hui Chen, Jie Yan, Xiangyu Zhang

机构 * StepFun University of Chinese Academy of Sciences(中国科学院大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Added links to GitHub and Hugging Face

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08171 2025-08-12 cs.SE cs.AI 57%

PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C

Pedro Orvalho, Marta Kwiatkowska

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 14 pages, 6 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18337 2025-08-06 cs.AI 57%

The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams

Peter Baumgartner, Lachlan McGinness

机构 * CSIRO/Data61 and Australian National University(CSIRO/Data61和澳大利亚国立大学) Australian National University and CSIRO/Data61(澳大利亚国立大学和CSIRO/Data61)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23694 2025-08-01 cs.MA cs.AI 57%

A survey of multi-agent geosimulation methodologies: from ABM to LLM

Virginia Padilla, Jacinto Dávila

机构 * Departamento de Ciencia y Tecnología, Universidad Nacional Experimental de Guayana(科学与技术系,圭亚那国家实验大学) CeSiMo, Facultad de Ingeniería, Universidad de los Andes(CeSiMo,工程学院,安第斯大学)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Comments 20 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22009 2025-07-30 cs.AI 57%

PHAX: A Structured Argumentation Framework for User-Centered Explainable AI in Public Health and Biomedical Sciences

Bahar İlgen, Akshat Dubey, Georges Hattab

机构 * Center for Artificial Intelligence in Public Health Research (ZKI-PH), Robert Koch-Institut, Nordufer 20, Berlin, 13353, Germany(公共健康人工智能研究中心(ZKI-PH)、罗伯特· Koch研究所) Department of Mathematics and Computer Science, Freie Universität Berlin, Arnimallee 14, Berlin, 14195, Germany(数学与计算机科学系、柏林自由大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21067 2025-07-30 cs.AI cs.CY cs.HC 57%

SynLang and Symbiotic Epistemology: A Manifesto for Conscious Human-AI Collaboration

Jan Kapusta

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 32 pages, 4 figures. Includes 2 Appendices containing SynLang v1.2.0 protocol specification, and formal BNF grammar

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18290 2025-07-25 cs.AI 57%

Foundations for Risk Assessment of AI in Protecting Fundamental Rights

Antonino Rotolo, Beatrice Ferrigno, Jose Miguel Angel Garcia Godinez, Claudio Novelli, Giovanni Sartor

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 24 pages, 1 figure. To be published in: The Philosophical Foundations of Information Technology Law. Oxford University Press, Oxford

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17134 2025-07-24 cs.MA cs.AI 57%

Resilient Multi-Agent Negotiation for Medical Supply Chains:Integrating LLMs and Blockchain for Transparent Coordination

Mariam ALMutairi, Hyungmin Kim

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 11 pages, 6 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13334 2025-07-22 cs.CL 57%

A Survey of Context Engineering for Large Language Models

Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, Chenlin Zhou, Jiayi Mao, Tianze Xia, Jiafeng Guo, Shenghua Liu

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学默塞德分校) The University of Queensland(昆士兰大学) Peking University(北京大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments ongoing work; 166 pages, 1411 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14335 2025-07-22 cs.AI 57%

ProofCompass: Enhancing Specialized Provers with LLM Guidance

Nicolas Wischermann, Claudio Mayrink Verdun, Gabriel Poesia, Francesco Noseda

机构 * Federal University of Rio de Janeiro(里约热内卢联邦大学) Harvard John A. Paulson School of Engineering and Applied Sciences(哈佛大学约翰·A·保罗森工程与应用科学学院) Stanford University(斯坦福大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 19 pages, 7 figures. Accepted at the 2nd AI for MATH Workshop at the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01789 2025-07-22 cs.SE cs.AI 57%

Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs

Shan Jiang, Chenguang Zhu, Sarfraz Khurshid

机构 * The University of Texas at Austin(德克萨斯大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11479 2025-07-16 cs.AI cs.GR cs.HC 57%

Perspective-Aware AI in Extended Reality

Daniel Platnick, Matti Gruener, Marjan Alirezaie, Kent Larson, Dava J. Newman, Hossein Rahnama

机构 * Flybits Labs(Flybits实验室) Creative Ai Hub(创意人工智能中心) Toronto Metropolitan University(多伦多 Metropolitan 大学) MIT Media Lab(MIT媒体实验室)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted to the International Conference on eXtended Reality (2025), 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11127 2025-07-16 cs.AI 57%

Defining neurosymbolic AI

Lennert De Smet, Luc De Raedt

机构 * Department of Computer Science, KU Leuven(计算机科学系,卢万大学) Örebro University(奥雷布罗大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09083 2025-07-15 cs.GT cs.AI 57%

Learning from Synthetic Labs: Language Models as Auction Participants

Anand Shah, Kehang Zhu, Yanchen Jiang, Jeffrey G. Wang, Arif K. Dayi, John J. Horton, David C. Parkes

机构 * Dropbox Inc.(Dropbox公司) Expected Parrot Inc.(Expected Parrot公司)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08362 2025-07-14 cs.LG 57%

Leveraging Machine Learning and Enhanced Parallelism Detection for BPMN Model Generation from Text

Phuong Nam Lê, Charlotte Schneider-Depré, Alexandre Goossens, Alexander Stevens, Aurélie Leribaux, Johannes De Smedt

机构 * Research Centre for Information Systems Engineering, KU Leuven, Leuven, Belgium(信息系统工程研究中心,鲁汶大学,比利时)

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02639 2025-07-04 cs.LG 57%

On Efficient Bayesian Exploration in Model-Based Reinforcement Learning

Alberto Caron, Chris Hicks, Vasilios Mavroudis

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23486 2025-07-04 cs.AI 57%

Autoformalization in the Era of Large Language Models: A Survey

Ke Weng, Lun Du, Sirui Li, Wangyue Lu, Haozhe Sun, Hengyu Liu, Tiancheng Zhang

机构 * Northeastern University(东北大学) Ant Research Institute, Ant Group(蚂蚁集团研究院) Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17506 2025-06-24 cs.CL cs.OS 57%

VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM

Lesheng Jin, Zhenyuan Ruan, Haohui Mai, Jingbo Shang

机构 * UC San Diego(加州大学圣地亚哥分校) MIT(麻省理工学院) CausalFlow Inc.(CausalFlow公司)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11906 2025-06-18 cs.CV cs.AI 57%

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension

Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 代码与定理证明 :chain-of-thought(abstract);分类 cs.AI

Comments This is the camera-ready version for ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13246 2025-06-17 cs.CR cs.AI cs.DC 57%

On Immutable Memory Systems for Artificial Agents: A Blockchain-Indexed Automata-Theoretic Framework Using ECDH-Keyed Merkle Chains

Craig Steven Wright

机构 * Dr Craig S. Wright University of Exeter Business School(埃克塞特大学商学院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 47 pages, includes formal automata specifications, cryptographic constructions, and epistemic architecture schema

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10984 2025-06-16 cs.SE cs.AI 57%

Application Modernization with LLMs: Addressing Core Challenges in Reliability, Security, and Quality

Ahilan Ayyachamy Nadar Ponnusamy

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02747 2025-06-12 cs.RO cs.AI cs.CR 57%

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

Hongwei Li, Yuheng Tang, Shiqi Wang, Wenbo Guo

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校) Meta

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07047 2025-06-10 cs.AI 57%

Mathesis: Towards Formal Theorem Proving from Natural Languages

Yu Xuejun, Jianyuan Zhong, Zijin Feng, Pengyi Zhai, Roozbeh Yousefzadeh, Wei Chong Ng, Haoxiong Liu, Ziyi Shou, Jing Xiong, Yudong Zhou, Claudia Beth Ong, Austen Jeremy Sugiarto, Yaoxi Zhang, Wai Ming Tai, Huan Cao, Dongcai Lu, Jiacheng Sun, Qiang Xu, Shen Xin, Zhenguo Li

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06034 2025-06-09 cs.CL 57%

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?

Zhitao He, Zongwei Lyu, Dazhong Chen, Dadi Guo, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05109 2025-06-06 cs.AI 57%

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning

Tennison Liu, Mihaela van der Schaar

机构 * DAMTP, University of Cambridge, Cambridge, UK(剑桥大学 DAMTP 实验室)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Comments Published as a conference paper at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23835 2025-06-02 cs.CL 57%

Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things

Ye Cheng, Minghui Xu, Yue Zhang, Kun Li, Hao Wu, Yechao Zhang, Shaoyong Guo, Wangjie Qiu, Dongxiao Yu, Xiuzhen Cheng

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) Institute of Artificial Intelligence, Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(北京航空航天大学人工智能研究院) Zhongguancun Laboratory, Beijing, China(中关村实验室,北京,中国)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04779 2025-05-30 cs.PL cs.AI cs.SE 57%

Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference

Thanh Le-Cong, Bach Le, Toby Murray

机构 * School of Computing and Information Systems(计算与信息系统学院) The University of Melbourne(墨尔本大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted to ACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16982 2025-05-23 cs.AI physics.med-ph 57%

Beyond Correlation: Towards Causal Large Language Model Agents in Biomedicine

Adib Bazgir, Amir Habibdoust Lafmajani, Yuwen Zhang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09338 2025-05-23 cs.CL 57%

Keys to Robust Edits: from Theoretical Insights to Practical Advances

Jianhao Yan, Futing Wang, Yun Luo, Yafu Li, Yue Zhang

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖先进研究院技术研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13664 2025-05-21 cs.CY cs.CL 57%

Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading

Ming Ding, Rasmus Kyng, Federico Solda, Weixuan Yuan

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏