arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1124 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1124 篇

2503.03840 2025-03-07 cs.CV cs.LG 57%

Decoupling the components of geometric understanding in Vision Language Models

Eliza Kosoy, Annya Dahmani, Andrew K. Lampinen, Iulia M. Comsa, Soojin Jeong, Ishita Dasgupta, Kelsey Allen

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.LG

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11113 2025-02-19 cs.CL 57%

Valuable Hallucinations: Realizable Non-realistic Propositions

Qiucheng Chen, Bo Wang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09913 2025-01-20 cs.AI 57%

Towards A Litmus Test for Common Sense

Hugo Latapie

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00900 2025-01-20 cs.SE cs.AI 57%

Large Process Models: A Vision for Business Process Management in the Age of Generative AI

Timotheus Kampik, Christian Warmuth, Adrian Rebmann, Ron Agam, Lukas N. P. Egger, Andreas Gerber, Johannes Hoffart, Jonas Kolk, Philipp Herzig, Gero Decker, Han van der Aa, Artem Polyvyanyy, Stefanie Rinderle-Ma, Ingo Weber, Matthias Weidlich

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Journal ref Künstl Intell (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01992 2025-01-07 cs.AI cs.LO cs.MA 57%

Disagree and Commit: Degrees of Argumentation-based Agreements

Timotheus Kampik, Juan Carlos Nieves

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments To appear eventually in the Autonomous Agents and Multi-Agent Systems journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01732 2025-01-06 cs.CR cs.AI cs.NI 57%

Combined Hyper-Extensible Extremely-Secured Zero-Trust CIAM-PAM architecture

Shivom Aggarwal, Shourya Mehra, Safeer Sathar

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01257 2025-01-06 cs.CL 57%

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, Yibo Miao, Yunlong Feng, Zekun Wang, Jian Yang, Zeyu Cui, Yang Fan, Yichang Zhang, Binyuan Hui, Junyang Lin

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21199 2025-01-03 cs.SE cs.CL 57%

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Zhaojian Yu, Yilun Zhao, Arman Cohan, Xiao-Ping Zhang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07636 2024-12-11 cs.CR cs.AI 57%

TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans

Md Omar Faruque, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15143 2024-11-26 cs.SE cs.AI cs.PL 57%

dafny-annotator: AI-Assisted Verification of Dafny Programs

Gabriel Poesia, Chloe Loughridge, Nada Amin

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13627 2024-11-22 cs.CR cs.AI cs.SC 57%

CryptoFormalEval: Integrating LLMs and Formal Verification for Automated Cryptographic Protocol Vulnerability Detection

Cristian Curaba, Denis D'Ambrosi, Alessandro Minisini, Natalia Pérez-Campanero Antolín

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05977 2024-11-11 cs.CL 57%

Mathematical Formalized Problem Solving and Theorem Proving in Different Fields in Lean 4

Xichen Tang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00828 2024-11-05 cs.CV cs.LG 57%

Dreaming Out Loud: A Self-Synthesis Approach For Training Vision-Language Models With Developmentally Plausible Data

Badr AlKhamissi, Yingtian Tang, Abdülkadir Gökce, Johannes Mehrer, Martin Schrimpf

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.LG

Comments Accepted to BabyLM Challenge at CoNLL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00188 2024-11-04 cs.AI cs.IR 57%

Building Multi-Agent Copilot towards Autonomous Agricultural Data Management and Analysis

Yu Pan, Jianxin Sun, Hongfeng Yu, Joe Luck, Geng Bai, Nipuna Chamara, Yufeng Ge, Tala Awada

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15880 2024-11-04 cs.PL cs.AI 57%

HYSYNTH: Context-Free LLM Approximation for Guiding Program Synthesis

Shraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick, Nadia Polikarpova

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23299 2024-11-01 cs.AR cs.AI 57%

FVEval: Understanding Language Model Capabilities in Formal Verification of Digital Hardware

Minwoo Kang, Mingjie Liu, Ghaith Bany Hamad, Syed Suhaib, Haoxing Ren

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21141 2024-10-29 cs.LG stat.ML 57%

LLM-initialized Differentiable Causal Discovery

Shiv Kampani, David Hidary, Constantijn van der Poel, Martin Ganahl, Brenda Miao

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12300 2024-10-15 cs.AI cs.GT 57%

Autoformalization of Game Descriptions using Large Language Models

Agnieszka Mensfelt, Kostas Stathis, Vince Trencsenyi

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments code: https://github.com/dicelab-rhul/game-formaliser

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01379 2024-10-14 cs.CL 57%

Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving

Xin Quan, Marco Valentino, Louise A. Dennis, André Freitas

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments Camera-ready for EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04397 2024-10-11 cs.CR cs.AI 57%

Towards Understanding and Enhancing Security of Proof-of-Training for DNN Model Ownership Verification

Yijia Chang, Hanrui Jiang, Chao Lin, Xinyi Huang, Jian Weng

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted by USENIX Security 2025 (Major Revision -> Accept)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03203 2024-10-07 cs.FL cs.AI 57%

TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts

Ruida Wang, Jipeng Zhang, Yizhen Jia, Rui Pan, Shizhe Diao, Renjie Pi, Tong Zhang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07453 2024-09-12 cs.AI cs.HC 57%

"My Grade is Wrong!": A Contestable AI Framework for Interactive Feedback in Evaluating Student Essays

Shengxin Hong, Chang Cai, Sixuan Du, Haiyue Feng, Siyuan Liu, Xiuyi Fan

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19238 2024-08-23 cs.AI 57%

Human-Aware Belief Revision: A Cognitively Inspired Framework for Explanation-Guided Revision of Human Models

Stylianos Loukas Vasileiou, William Yeoh

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09939 2024-08-23 cs.AI 57%

A Survey on Deep Learning for Theorem Proving

Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, Xujie Si

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00147 2024-08-02 cs.AI cs.LO 57%

Formal Ethical Obligations in Reinforcement Learning Agents: Verification and Policy Updates

Colin Shea-Blymyer, Houssam Abbas

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12188 2024-07-23 cs.AI stat.AP 57%

ChatGPT and post-test probability

Samuel J. Weisenthal

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 138 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07408 2024-06-06 cs.CR cs.AI 57%

Large Language Models are Few-shot Generators: Proposing Hybrid Prompt Algorithm To Generate Webshell Escape Samples

Mingrui Ma, Lansheng Han, Chunjie Zhou

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 21 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14333 2024-05-24 cs.AI 57%

DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, Xiaodan Liang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08848 2024-05-16 cs.SE cs.AI 57%

Automated Repair of AI Code with Large Language Models and Formal Verification

Yiannis Charalambous, Edoardo Manino, Lucas C. Cordeiro

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00011 2024-04-02 cs.HC cs.CL 57%

A novel interface for adversarial trivia question-writing

Jason Liu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments 17 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏