Mind's Eye: Grounded Language Model Reasoning through Simulation
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to COLING 2022
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 18 pages, 4 figures, 10 tables, accepted in Findings of the Association for Computational Linguistics 2022
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted for publication at EMNLP 2021, 11 pages, 5 tables, 4 figures
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments EMNLP 2021. Code and data are available at https://github.com/WadeYin9712/GD-VCR
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments We have updated this paper with considerable modifications
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments ACL 2021 Findings
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments To appear at ACL2021
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments In Proceedings of Findings of the Association for Computational Linguistics: ACL 2021 (ACL-Findings). Contains 16 pages, 14 figures and 11 tables
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments EACL 2021 (14 pages, 2 figures, 10 tables)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments CVPR 2021 paper. Supplementary: http://wellyzhang.github.io/attach/cvpr21zhang_acre_supp.pdf Project: http://wellyzhang.github.io/project/acre.html
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
Comments 22 pages, NeurIPS 2020
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2020 Findings. Add one more human reference for each test example: Table 1,3 & Figure 4 & Section 3.3, 3.4 are updated. Project page: https://inklab.usc.edu/CommonGen/
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
Comments ACL 2020
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 11 pages
Journal ref AKBC 2020
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 9pages
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2019
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
Comments 10 pages, 7 tables, 2 figures
为策略制定者搭建框架:霍特林空间市场中与架构相关的推理干预
机构 * California Institute of Technology(加州理工学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 研究结构化推理干预对大语言模型战略经济推理的影响及与模型架构的关系,以霍特林模型评估GPT - 4.1 - mini和GPT - 5 - mini,发现支架类型与模型架构有交叉交互作用,对抗性测试有损害,还存在陈述性 - 程序性差距。
Comments 26 pages (11 main + 15 appendix), 6 figures, 4 tables. Accepted at the ICLR 2026 Workshop on LLM Reasoning
面向旅游的推理大语言模型:基于领域特定知识图谱
机构 * Expedia Group(Expedia集团)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 提出一种模块化流水线,利用专家构建的知识图谱生成多跳问答对,微调大语言模型以提升旅游领域推理的准确性和校准能力,在基准上达到82.4%精确匹配。
Comments Accepted to the Uncertainty Reasoning and Quantification in Decision Making (UDM) Workshop, KDD 2026 (To be presented in August 2026)
ComBench: 奥林匹克级组合数学中严格证明推理与构造实现的基准测试
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Peking University(北京大学) ; Shanghai Jiao Tong University(上海交通大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 提出ComBench基准,包含100道奥林匹克级组合问题,分分析和构造两类,通过评分与验证评估大模型推理能力,发现最强模型准确率仅65.4%,且证明推理与构造实现能力存在差异。
Comments 39 pages, 6 figures, 26 tables. Project page: https://simplified-reasoning.github.io/ComBench/docs/
视觉-表格问答:面向表格图像推理的开放领域基准
机构 * École de Technologie Supérieure(埃克塞技术学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 本文提出Visual-TableQA,一个大规模开放领域多模态数据集,用于评估和提升视觉推理能力,包含2500个LaTeX渲染表格和6000对推理密集型问答对,通过多模型协作生成,验证了模型在外部基准上的鲁棒性。
Comments Accepted at the First Workshop on Foundations of Reasoning in Language Models, NeurIPS 2025. Available at: https://openreview.net/forum?id=fvJRsGwhPf
基于LLM的多模态推理用于加密流量解释:一个基准
机构 * School of Microelectronics and Communication Engineering, Chongqing University(重庆大学微电子与通信工程学院) ; School of Data Science, Lingnan University(岭南大学数据科学学院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 本文提出BGTD基准和mmTraffic框架,通过结合原始字节与结构化注释,实现可解释的加密流量解释,生成高保真的人可读报告,同时保持高分类准确率。
Comments Project page \url{https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark}
并非其他:一种区分推理与记忆的通用技术,用于多选LLM评估基准
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL
AI总结 本文提出一种通用方法,通过改变数学问题的数值来区分LLM的推理能力与记忆能力,评估了多个模型在公开和私有数据集上的表现,发现模型在该方法下准确率显著下降,揭示了记忆在当前LLM回答中的重要作用。
Journal ref "On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks," in IEEE Access, vol. 14, pp. 9384-9393, 2026