An Overview Of Temporal Commonsense Reasoning and Acquisition
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 27 pages, 7 figures, 6 tables
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 27 pages, 7 figures, 6 tables
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 11 pages, 6 figures, 5 tables, camera ready version of SIGKDD 2023
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments ECCV 2022, 51 pages, 23 figures, 4 tables
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Corrected typos. Accepted to NeurIPS 2021, 27 pages, 18 figures. Data and code are available at https://iconqa.github.io
专题命中 推理评测 :reasoning(title,abstract);planning(abstract)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2021. Project page: http://ptr.csail.mit.edu/
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2020
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, published at ICLR 2021
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments AAAI 2020
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 7 pages, 3 figures, presented at 2017 ICML Workshop on Human Interpretability in Machine Learning (WHI 2017), Sydney, NSW, Australia
大型语言模型的推理结构
机构 * ETH Zurich, Switzerland(苏黎世联邦理工学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
AI总结 针对大型推理模型评估中隐藏不同推理结构的问题,提出基于逻辑谜题的基准测试和将非结构化轨迹转化为可验证推理图的方法,并定义推理效率指标,以量化分析推理拓扑结构。
Comments Accepted at ICML 2026 and presented at the ICLR 2026 workshop on LLM reasoning
后训练推理数据入门:我们对其运作机制的了解
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本文综述了后训练推理数据的类型、效用、构建方法和扩展规律,为未来推理数据发布和后训练方案提供归因框架。
Comments 22 pages. Project Repository: https://github.com/RenBing-Sumeru/Awesome-LLM-Reasoning-Data
显式推理使评判更可靠:对准确性、效率和鲁棒性的系统研究
机构 * Arizona State University(亚利桑那州立大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本文通过对比显式推理与非显式推理模型在RewardBench任务中的表现,发现显式推理模型在准确性、效率和鲁棒性上均优于非显式模型,且在多语言环境下也表现出优势。
Comments Accepted in 2025 NeurIPS Foundations of Reasoning in Language Models Workshop
推理如何从训练数据中演变:使用国际象棋的实证研究
机构 * Department of Computer Science and Engineering, University of California San Diego, San Diego, United States(计算机科学与工程系,加州大学圣地亚哥分校,圣地亚哥,美国)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
AI总结 研究语言模型推理从监督微调到强化学习的演变,发现直接预测最佳走法的微调能提升性能,但强化学习阶段导致不一致推理,而多步轨迹训练则能获得更稳定的推理。
Comments Accepted at ICML 2026. An earlier version appeared at the NeurIPS 2025 Foundations of Reasoning in Language Models (FoRLM) Workshop (Oral)
缓解语言模型中的捷径推理:一种梯度感知的训练方法
机构 * Arizona State University(亚利桑那州立大学) ; Clemson University(克莱姆森大学) ; University of Kansas(堪萨斯大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出SART框架,通过梯度感知技术检测并缓解语言模型中的捷径推理,实验显示在受控推理基准上提升准确率和鲁棒性。
Comments 12 pages, 2 figures. Preprint. Experiments on synthetic reasoning benchmarks. Code available
我了解我所不知道的:用于多证据概率推理的潜在后验因子模型
机构 * Epalea
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
AI总结 本文提出LPF模型,通过将变分自编码器的潜在后验转换为软似然因子,实现对无结构证据的可 tractable 概率推理,同时保持校准的不确定性估计。
Comments 202 pages, 52 figures, 105 tables. Comprehensive presentation of the Latent Posterior Factors (LPF) framework for multi-evidence probabilistic reasoning, including theoretical analysis, algorithmic design, and extensive empirical evaluation across synthetic and real-world benchmarks
OfficeQA Pro:一个企业级端到端 grounded 推理基准测试
机构 * Databricks AI Research(Databricks人工智能研究)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 OfficeQA Pro 是一个企业级基准测试,评估 AI 代理在大规模文档语料库上进行端到端 grounded 推理的能力,发现结构化文档表示可显著提升性能。
Comments 24 pages, 16 figures. Introduces the OfficeQA Pro benchmark for grounded reasoning over enterprise documents
FlagEval Findings Report: 对大推理模型在自动可验证文本和视觉问题上的初步评估
机构 * BAAI FlagEval Team(百度人工智能研究院FlagEval团队) ; State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室) ; Peking University(北京大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI总结 FlagEval报告对大推理模型在自动可验证文本和视觉问题上的初步评估进行了研究,提出了ROME基准以测试视觉线索推理能力。
Comments Project homepage: https://flageval-baai.github.io/LRM-Eval/ This work will also be presented at NeurIPS 2025 Workshop on Foundations of Reasoning in Language Models (FoRLM); update with trials on Gemini 3 Pro
机构 * Personalization Team, Walmart Global Tech(Walmart全球科技个性化团队)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning
机构 * Independent Researcher(独立研究者)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted by 39th NeurIPS - Foundations of Reasoning in Language Models
机构 * Department of Computer Science and Engineering, University of Michigan, Ann Arbor, USA(计算机科学与工程系,密歇根大学,安娜堡,美国)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments Title updated to "Token-Efficient RL for LLM Reasoning" to better reflect algorithmic focus. Revised abstract, intro, and conclusion. Paper shortened and typos fixed
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 39 pages. Accepted at the NeurIPS 2024 Workshops on Mathematical Reasoning and AI and Open-World Agents
重新审视Agentic-SQL:面向LLM文本转SQL的基于自主性的分类与实证基准分析
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.AI
AI总结 本研究针对LLM文本转SQL领域,构建基于自主性的分类框架,通过Spider案例研究分析8B开源模型与DeepSeek-V3等基线的表现,发布工具链用于构建可追溯的基准排行榜。
LLM编排何时划算?关于准确率、成本与任务难度的对照评估
机构 * Zuse Institute Berlin(柏林Zuse研究所) ; TU Berlin(柏林工业大学) ; Weizenbaum Institute Berlin(柏林魏茨曼研究所)
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.AI
AI总结 该研究对照评估了Self-Refine等三种LLM编排方法与基线方法在5种骨干模型、3个领域的表现,发现编排收益依赖模型,需权衡准确率提升与额外推理成本。
视觉语言模型的持续学习:超越遗忘的综述与分类
专题命中 推理评测 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.LG
AI总结 本文综述了视觉语言模型的持续学习挑战,提出四种核心范式以解决跨模态特征漂移和灾难性遗忘问题,强调零样本学习和智能体生态系统的发展。