arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3024 信号源:cs.CL, cs.AI, cs.LG

1. 逻辑推理 3024 篇

2512.22258 2025-12-30 cs.AI cs.LG cs.LO cs.SC 62%

Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method

逻辑草图提示(LSP):一种确定性和可解释的提示方法

Satvik Tripathi

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 逻辑草图提示(LSP)通过引入类型变量、确定性条件评估器和规则验证器,提升了大语言模型在需要严格规则遵守、确定性和可审计性任务中的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17923 2025-12-30 q-fin.ST cs.AI cs.LG 62%

Inferring Latent Market Forces: Evaluating LLM Detection of Gamma Exposure Patterns via Obfuscation Testing

推断潜在市场力量:通过混淆测试评估大语言模型检测伽马暴露模式的能力

Christopher Regan, Ying Xie

机构 * Department of Computer Science Kennesaw State University Marietta, GA, Cobb(计算机科学系 凯尼恩州立大学 马里埃塔, 乔治亚州, 理查兹) Department of Information Technology Professor of Information Technology Kennesaw State University Marietta, GA(信息科技系 信息科技教授 凯尼恩州立大学 马里埃塔, 乔治亚州)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过混淆测试验证大语言模型能否通过因果推理识别结构性市场模式,发现其在无偏提示下检测伽马暴露模式的准确率为71.5%,展示了LLMs在金融机制识别方面的潜力。

Comments 10 pages, 8 figures. Accepted at IEEE Big Data 2025. Extended journal version in preparation. ISBN: 979-8-3315-9447-3/25. Page numbers: 7226-7235

Journal ref 2025 IEEE International Conference on Big Data (Big Data)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17920 2025-12-23 cs.CL cs.AI 62%

Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression

分离约束合规性与语义准确性:一种新的基准,用于在压缩下评估指令遵循

Rahul Baxi

机构 * Independent Researcher(独立研究者)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CDCT基准,揭示LLMs在压缩下约束合规性与语义准确性之间的矛盾,发现中等压缩时约束违规主要由RLHF训练的有用性行为导致。

Comments 19 pages, 9 figures; currently under peer review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17308 2025-12-22 cs.AI cs.CL 62%

Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation

大型语言模型作为宝可梦战斗代理:战略玩法与内容生成

Daksh Jain, Aarya Jain, Ashutosh Desai, Avyakt Verma, Ishan Bhanuka, Pratik Narang, Dhruv Kumar

机构 * Birla Institute of Technology and Science, Pilani, India(比拉理工学院和科学学院,比兰)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了大型语言模型在宝可梦战斗中的战略决策能力及内容生成能力,展示了其作为动态游戏对手和设计者的潜力。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15298 2025-12-18 cs.AI cs.CL cs.CY 62%

ChatGPT and Gemini participated in the Korean College Scholastic Ability Test -- Earth Science I

ChatGPT 和 Gemini 参与韩国大学入学考试——地球科学I

Seok-Hyun Ga, Chun-Yen Chang

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本研究分析了ChatGPT和Gemini在地球科学I考试中的多模态推理能力,发现模型在感知与认知之间存在差距,揭示了AI在科学评估中的局限性。

Comments 23 pages, 9 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12012 2025-12-17 cs.CV cs.AI cs.CL cs.RO 62%

Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus

语义驱动:通过开放词汇锚定和神经符号视觉语言共识民主化长尾数据整理

Antonio Guillen-Perez

机构 * Independent Researcher(独立研究者)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 Semantic-Drive通过开放词汇锚定和神经符号视觉语言共识,提升自动驾驶中长尾数据整理的效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18054 2025-12-17 cs.IR cs.AI cs.LG 62%

A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem

基于知识图谱的检索增强生成框架用于设施布局问题中的算法选择

Nikhil N S, Bilal Muhammed, Soban Babu Beemaraj, Amol Dilip Joshi

机构 * Indian Institute of Science(印度科学学院) TCS Research and Innovation(塔塔咨询公司研究与创新)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于知识图谱的检索增强生成框架,用于自动推荐设施布局问题中的算法选择,通过多维检索机制和大语言模型实现数据驱动的算法推荐。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10121 2025-12-12 cs.CL cs.AI cs.CY q-fin.GN 62%

Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing

流程即一切:通过高熵信息觅食和对抗性节奏控制摆脱'统计平滑陷阱

Zhongjie Jiang

机构 * Zhongjie Jiang(江宗杰)

专题命中 逻辑推理 :planning(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出DeepNews框架,通过高熵信息觅食和对抗性节奏控制,解决LLMs在长文本生成中的幻觉、逻辑一致性和个性化表达难题,实验显示其在金融报道中表现优异。

Comments 22 pages, 8 figures. Includes an ecological validity blind test where the Agentic Workflow achieved a 25% acceptance rate in top-tier media, decisively outperforming the SOTA Zero-shot baseline (0%). Features the DNFO-v5 ontology

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15843 2025-12-11 cs.CL cs.AI 62%

Enhanced Sentiment Interpretation via a Lexicon-Fuzzy-Transformer Framework

通过词典-模糊-变换器框架提升情感解释

Shayan Rokhva, Mousa Alizadeh, Maryam Abdollahi Shamami

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种结合规则启发式、上下文深度学习和模糊逻辑的混合框架,用于提升领域特定文本中情感极性和强度的识别精度。

Comments The manuscript was uploaded in error and is scientifically invalid. It is an incomplete draft with major flaws. Co-authors were not aware of or consenting to this submission and do not endorse it

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00409 2025-12-05 cs.CL cs.AI 62%

Semantic Mastery: Enhancing LLMs with Advanced Natural Language Understanding

语义 mastery:通过高级自然语言理解增强大语言模型

Mohanakrishnan Hariharan

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文通过高级自然语言理解技术提升大语言模型的语义理解和推理能力,探讨了知识图谱、检索增强生成和对比学习等方法在解决复杂NLP任务中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02987 2025-12-03 cs.CL cs.AI 62%

Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic

针对逻辑翻译的微调大型语言模型:通过Lang2Logic减少幻觉

Muyu Pan, Dheeraj Kodakandla, Mahfuza Farooque

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种通过微调大型语言模型来减少逻辑翻译中幻觉的框架,利用自定义语法和符号计算库生成可靠的合取范式。

Comments IEEE ISNCC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17638 2025-11-25 cs.LG cs.AI 62%

Model-to-Model Knowledge Transmission (M2KT): A Data-Free Framework for Cross-Model Understanding Transfer

模型到模型知识传输(M2KT):一种无数据的跨模型理解迁移框架

Pratham Sorte

机构 * Department of Computer Science(计算机科学系) Engineering MIT-World Peace University, Pune, India(工程学院 MIT-世界和平大学 印度邦普尼)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 M2KT提出了一种无数据的跨模型知识传输方法,通过概念空间交换知识包,实现高效的知识迁移和模型自我改进。

Comments 8 pages including figures, prepared in IEEE conference style. Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14334 2025-11-20 cs.AI cs.LG 62%

When Words Change the Model: Sensitivity of LLMs for Constraint Programming Modelling

Alessio Pellegrino, Jacopo Mauro

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17053 2025-11-18 cs.LG cs.AI cs.CR 62%

DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately

Huiwen Wu, Deyi Zhang, Xiaohan Li, Xiaogang Xu, Jiafei Wu, Zhe Liu

机构 * Zhejiang Laboratory(浙江实验室)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09956 2025-11-17 cs.LG cs.AI cs.CV cs.ET 62%

DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey

Yu Qiao, Phuong-Nam Tran, Ji Su Yoon, Loc X. Nguyen, Eui-Nam Huh, Dusit Niyato, Choong Seon Hong

机构 * Kyung Hee University(韩国庆熙大学) Nanyang Technological University(南洋理工大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 45 pages, 12 figures

Journal ref ACM Computing Surveys, Nov. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18890 2025-11-17 cs.AI cs.LG cs.NE 62%

CoEvo: Continual Evolution of Symbolic Solutions Using Large Language Models

Ping Guo, Qingfu Zhang, Xi Lin

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Camera ready version for AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18573 2025-11-14 cs.CL cs.AI 62%

FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models

Radu Marinescu, Debarun Bhattacharjya, Junkyu Lee, Tigran Tchrakian, Javier Carnerero Cano, Yufang Hou, Elizabeth Daly, Alessandra Pascale

机构 * IBM Research(IBM研究院) IT:U - Interdisciplinary Transformation University Austria(interdisciplinary Transformation University Austria)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16335 2025-11-11 cs.AI cs.CL 62%

Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach

Albert Sadowski, Jarosław A. Chudziak

机构 * Faculty of Electronics and Information Technology(电子与信息技术学院) Warsaw University of Technology(华沙技术大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 29th International Conference on Knowledge-Based and Intelligent Information \& Engineering Systems (KES 2025)

Journal ref Procedia Computer Science 270 (2025) 2166-2175

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04566 2025-11-04 cs.RO cs.AI cs.CV cs.LG 62%

Knolling Bot: Teaching Robots the Human Notion of Tidiness

Yuhang Hu, Judah Goldfeder, Zhizhuo Zhang, Xinyue Zhu, Ruibo Liu, Philippe Wyder, Jiong Lin, Hod Lipson

机构 * Columbia University(哥伦比亚大学)

专题命中 逻辑推理 :planning(abstract);分类 cs.AI、cs.LG

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10455 2025-10-31 cs.CL cs.AI 62%

IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction

Kaiyu He, Mian Zhang, Shuo Yan, Peilin Wu, Zhiyu Zoey Chen

机构 * Department of Computer Science University of Texas at Dallas(计算机科学系德克萨斯大学达拉斯分校)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21204 2025-10-31 cs.AI cs.CL 62%

Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding

Vanessa Figueiredo

机构 * Department of Computer Science University of Regina(计算机科学系 皇家大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17022 2025-10-30 cs.LG cs.AI 62%

Curiosity-driven RL for symbolic equation solving

Kevin P. O'Keeffe

机构 * Starling Research Institute(星灵研究机构)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments Accepted at the NeurIPS 2025 MATH-AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25445 2025-10-30 cs.AI cs.LG 62%

Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions

Mohamad Abou Ali, Fadi Dornaika

机构 * University of the Basque Country(巴斯克大学) Lebanese International University (LIU)(黎巴嫩国际大学) The International University of Beirut(贝鲁特国际大学) IKERBASQUE

专题命中 逻辑推理 :planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23682 2025-10-29 cs.LG cs.AI cs.LO cs.SE 62%

Beyond Prompt Engineering: Neuro-Symbolic-Causal Architecture for Robust Multi-Objective AI Agents

Gokturk Aytug Akarlar

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 35 pages, 15 figures, 2 tables. Keywords: Large Language Models, Autonomous Agents, Neuro-Symbolic AI, Causal Inference, Formal Verification, Multi-Objective Optimization. Open-source code and interactive demo available

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20643 2025-10-22 cs.CL cs.AI 62%

Ontology-Enhanced Knowledge Graph Completion using Large Language Models

Wenbin Guo, Xin Wang, Jiaoyan Chen, Zhao Li, Zirui Chen

机构 * Tianjin University, College of Intelligence and Computing(天津大学智能与计算学院) University of Manchester, Department of Computer Science(曼彻斯特大学计算机科学系)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17910 2025-10-22 cs.CY cs.AI cs.CL 62%

Interpretability Framework for LLMs in Undergraduate Calculus

Sagnik Dakshit, Sushmita Sinha Roy

机构 * University of Texas at Tyler(德克萨斯理工大学) Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16381 2025-10-21 cs.CL cs.AI 62%

ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents

David Peer, Sebastian Stabinger

机构 * Otera Austria(奥特拉奥地利)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11943 2025-10-21 cs.AI cs.LG cs.LO cs.MA 62%

Agentic System with Modal Logic for Autonomous Diagnostics

Antonin Sulc, Thorsten Hellert

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12269 2025-10-17 cs.AI cs.LG cs.NE cs.PL stat.ML 62%

Tensor Logic: The Language of AI

Pedro Domingos

机构 * University of Washington(华盛顿大学)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20921 2025-10-17 cs.LG cs.AI cs.CE 62%

LLM-guided Chemical Process Optimization with a Multi-Agent Approach

Tong Zeng, Srivathsan Badrinarayanan, Janghoon Ock, Cheng-Kai Lai, Amir Barati Farimani

机构 * Department of Chemical Engineering, Carnegie Mellon University(化学工程系,卡内基梅隆大学) Department of Mechanical Engineering, Carnegie Mellon University(机械工程系,卡内基梅隆大学) Department of Biomedical Engineering, Carnegie Mellon University(生物医学工程系,卡内基梅隆大学) Machine Learning Department, Carnegie Mellon University(机器学习系,卡内基梅隆大学) Department of Chemical and Biomolecular Engineering, University of Nebraska--Lincoln(化学与生物分子工程系,内布拉斯加大学林肯分校)

专题命中 逻辑推理 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 16 pages (main manuscript without references), 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏