arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1124 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1124 篇

2505.23486 2025-07-04 cs.AI 57%

Autoformalization in the Era of Large Language Models: A Survey

Ke Weng, Lun Du, Sirui Li, Wangyue Lu, Haozhe Sun, Hengyu Liu, Tiancheng Zhang

机构 * Northeastern University(东北大学) Ant Research Institute, Ant Group(蚂蚁集团研究院) Department of Computer Science, Aalborg University(奥尔堡大学计算机科学系)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17506 2025-06-24 cs.CL cs.OS 57%

VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM

Lesheng Jin, Zhenyuan Ruan, Haohui Mai, Jingbo Shang

机构 * UC San Diego(加州大学圣地亚哥分校) MIT(麻省理工学院) CausalFlow Inc.(CausalFlow公司)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11906 2025-06-18 cs.CV cs.AI 57%

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension

Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 代码与定理证明 :chain-of-thought(abstract);分类 cs.AI

Comments This is the camera-ready version for ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13246 2025-06-17 cs.CR cs.AI cs.DC 57%

On Immutable Memory Systems for Artificial Agents: A Blockchain-Indexed Automata-Theoretic Framework Using ECDH-Keyed Merkle Chains

Craig Steven Wright

机构 * Dr Craig S. Wright University of Exeter Business School(埃克塞特大学商学院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 47 pages, includes formal automata specifications, cryptographic constructions, and epistemic architecture schema

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10984 2025-06-16 cs.SE cs.AI 57%

Application Modernization with LLMs: Addressing Core Challenges in Reliability, Security, and Quality

Ahilan Ayyachamy Nadar Ponnusamy

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02747 2025-06-12 cs.RO cs.AI cs.CR 57%

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

Hongwei Li, Yuheng Tang, Shiqi Wang, Wenbo Guo

机构 * University of California, Santa Barbara(加州大学圣芭芭拉分校) Meta

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07047 2025-06-10 cs.AI 57%

Mathesis: Towards Formal Theorem Proving from Natural Languages

Yu Xuejun, Jianyuan Zhong, Zijin Feng, Pengyi Zhai, Roozbeh Yousefzadeh, Wei Chong Ng, Haoxiong Liu, Ziyi Shou, Jing Xiong, Yudong Zhou, Claudia Beth Ong, Austen Jeremy Sugiarto, Yaoxi Zhang, Wai Ming Tai, Huan Cao, Dongcai Lu, Jiacheng Sun, Qiang Xu, Shen Xin, Zhenguo Li

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06034 2025-06-09 cs.CL 57%

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?

Zhitao He, Zongwei Lyu, Dazhong Chen, Dadi Guo, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05109 2025-06-06 cs.AI 57%

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning

Tennison Liu, Mihaela van der Schaar

机构 * DAMTP, University of Cambridge, Cambridge, UK(剑桥大学 DAMTP 实验室)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Comments Published as a conference paper at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23835 2025-06-02 cs.CL 57%

Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things

Ye Cheng, Minghui Xu, Yue Zhang, Kun Li, Hao Wu, Yechao Zhang, Shaoyong Guo, Wangjie Qiu, Dongxiao Yu, Xiuzhen Cheng

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) Institute of Artificial Intelligence, Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(北京航空航天大学人工智能研究院) Zhongguancun Laboratory, Beijing, China(中关村实验室,北京,中国)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04779 2025-05-30 cs.PL cs.AI cs.SE 57%

Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference

Thanh Le-Cong, Bach Le, Toby Murray

机构 * School of Computing and Information Systems(计算与信息系统学院) The University of Melbourne(墨尔本大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted to ACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16982 2025-05-23 cs.AI physics.med-ph 57%

Beyond Correlation: Towards Causal Large Language Model Agents in Biomedicine

Adib Bazgir, Amir Habibdoust Lafmajani, Yuwen Zhang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09338 2025-05-23 cs.CL 57%

Keys to Robust Edits: from Theoretical Insights to Practical Advances

Jianhao Yan, Futing Wang, Yun Luo, Yafu Li, Yue Zhang

机构 * Zhejiang University(浙江大学) School of Engineering, Westlake University(西湖大学工程学院) Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖先进研究院技术研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13664 2025-05-21 cs.CY cs.CL 57%

Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading

Ming Ding, Rasmus Kyng, Federico Solda, Weixuan Yuan

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10982 2025-05-19 cs.AI 57%

Facets in Argumentation: A Formal Approach to Argument Significance

Johannes Fichte, Nicolas Fröhlich, Markus Hecher, Victor Lagerkvist, Yasir Mahmood, Arne Meier, Jonathan Persson

机构 * Department of Computer and Information Science, Linköping University, Sweden(林雪平大学计算机与信息科学系) Leibniz Universität Hannover, Germany(汉诺威莱布尼茨大学) Univ. Artois, CNRS, France(阿诺伊大学) CSAIL, Massachusetts Institute of Technology, USA(麻省理工学院计算机科学与人工智能实验室) DICE group, Paderborn University, Germany(波德生大学DICE小组)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10962 2025-05-19 cs.AI 57%

MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation

Zhenwen Liang, Linfeng Song, Yang Li, Tao Yang, Feng Zhang, Haitao Mi, Dong Yu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10387 2025-05-16 cs.MA cs.AI cs.CC 57%

Multi-Agent Path Finding For Large Agents Is Intractable

Artem Agafonov, Konstantin Yakovlev

机构 * HSE University(莫斯科国立高等经济大学) FRC CSC RAS(俄罗斯科学院应用系统分析研究所) AIRI(人工智能研究所)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08532 2025-05-14 cs.SI cs.AI 57%

The Truth Becomes Clearer Through Debate! Multi-Agent Systems with Large Language Models Unmask Fake News

Yuhan Liu, Yuxuan Liu, Xiaoqing Zhang, Xiuying Chen, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(仁爱大学人工智能学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19112 2025-05-13 cs.AI cs.CY cs.LO 57%

Logical Modalities within the European AI Act: An Analysis

Lara Lawniczak, Christoph Benzmüller

机构 * University of Bamberg(巴姆堡大学) Freie Universität Berlin(柏林自由大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Extended preprint of paper accepted for ICAIL 2025; 15 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03171 2025-05-07 cs.AI 57%

CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics

Junqi Liu, Xiaohan Lin, Jonas Bayer, Yael Dillies, Weijie Jiang, Xiaodan Liang, Roman Soletskyi, Haiming Wang, Yunzhou Xie, Beibei Xiong, Zhengfeng Yang, Jujian Zhang, Lihong Zhi, Jia Li, Zhengying Liu

机构 * Academy of Mathematics and Systems Science, University of Chinese Academy of Sciences(中国科学院数学与系统科学研究院) Sun Yat-sen University(中山大学) University of Cambridge(剑桥大学) East China Normal University(华东师范大学) Imperial College London(伦敦帝国理工学院) Stockholm Universitet(斯德哥尔摩大学) Numina Moonshot AI

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00841 2025-05-05 cs.CR cs.AI 57%

From Texts to Shields: Convergence of Large Language Models and Cybersecurity

Tao Li, Ya-Ting Yang, Yunian Pan, Quanyan Zhu

机构 * New York University(纽约大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11901 2025-04-15 cs.CL cs.PL cs.SE 57%

Building A Proof-Oriented Programmer That Is 64% Better Than GPT-4o Under Data Scarcity

Dylan Zhang, Justin Wang, Tianran Sun

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09067 2025-04-08 cs.CV cs.CL q-bio.NC 57%

Interpreting the structure of multi-object representations in vision encoders

Tarun Khajuria, Braian Olmiro Dias, Marharyta Domnich, Jaan Aru

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06967 2025-04-03 cs.AI cs.LO 57%

Data quality dimensions for fair AI

Camilla Quaresmini, Giuseppe Primiero

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00063 2025-04-02 cs.AI math.LO 57%

The Axiom-Based Atlas: A Structural Mapping of Theorems via Foundational Proof Vectors

Harim Yoo

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20561 2025-03-27 cs.LG stat.ML 57%

A Theoretical Framework for Prompt Engineering: Approximating Smooth Functions with Transformer Prompts

Ryumei Nakada, Wenlong Ji, Tianxi Cai, James Zou, Linjun Zhang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.LG

Comments 55 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18899 2025-03-25 cs.AI cs.CR 57%

Statistical Proof of Execution (SPEX)

Michele Dallachiesa, Antonio Pitasi, David Pinger, Josh Goodbody, Luis Vaello

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10694 2025-03-17 cs.CL 57%

Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, Travis Zack

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07172 2025-03-11 cs.AI cs.LO cs.SE 57%

Lawful and Accountable Personal Data Processing with GDPR-based Access and Usage Control in Distributed Systems

L. Thomas van Binsbergen, Marten C. Steketee, Milen G. Kebede, Heleen L. Janssen, Tom M. van Engers

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Submitted for review to the Journal of AI and Law, 49 pages (including)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03911 2025-03-07 cs.RO cs.LG cs.SY eess.SY 57%

Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis

Ahmad Hafez, Alireza Naderi Akhormeh, Amr Hegazy, Amr Alanwar

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏