arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5804 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5804 篇

2305.03796 2023-05-09 cs.CL 74%

Transformer Working Memory Enables Regular Language Reasoning and Natural Language Length Extrapolation

Ta-Chung Chi, Ting-Han Fan, Alexander I. Rudnicky, Peter J. Ramadge

专题命中 其他推理 :reasoning(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13958 2023-04-28 cs.CL cs.CY cs.SI 74%

Learning and Reasoning Multifaceted and Longitudinal Data for Poverty Estimates and Livelihood Capabilities of Lagged Regions in Rural India

Atharva Kulkarni, Raya Das, Ravi S. Srivastava, Tanmoy Chakraborty

专题命中 其他推理 :reasoning(title);分类 cs.CL

Comments Accepted to IJCAI 2023 Main Conference (AI for Social Good Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17012 2023-04-18 physics.ed-ph cs.AI 74%

Advances in apparent conceptual physics reasoning in GPT-4

Colin G. West

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments 5 pages, one figure one table. arXiv admin note: text overlap with arXiv:2303.01067 (longer, prior version of this project)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11658 2023-03-27 cs.CL 74%

Penguins Don't Fly: Reasoning about Generics through Instantiations and Exceptions

Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathleen McKeown, Doug Downey, Yejin Choi

专题命中 其他推理 :reasoning(title);分类 cs.CL

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05392 2022-11-11 cs.CL 74%

EvEntS ReaLM: Event Reasoning of Entity States via Language Models

Evangelia Spiliopoulou, Artidoro Pagnoni, Yonatan Bisk, Eduard Hovy

专题命中 其他推理 :reasoning(title);分类 cs.CL

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04634 2021-09-13 cs.LO cs.AI cs.SC 74%

Knowledge-Assisted Reasoning of Model-Augmented System Requirements with Event Calculus and Goal-Directed Answer Set Programming

Brendan Hall, Sarat Chandra Varanasi, Jan Fiedor, Joaquín Arias, Kinjal Basu, Fang Li, Devesh Bhatt, Kevin Driscoll, Elmer Salazar, Gopal Gupta

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments In Proceedings HCVS 2021, arXiv:2109.03988

Journal ref EPTCS 344, 2021, pp. 79-90

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.11743 2020-07-24 cs.MA cs.AI cs.LO 74%

Adaptable and Verifiable BDI Reasoning

Peter Stringer, Rafael C. Cardoso, Xiaowei Huang, Louise A. Dennis

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments In Proceedings AREA 2020, arXiv:2007.11260

Journal ref EPTCS 319, 2020, pp. 117-125

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.07489 2020-01-15 cs.AI 74%

Epistemic Graphs for Representing and Reasoning with Positive and Negative Influences of Arguments

Anthony Hunter, Sylwia Polberg, Matthias Thimm

专题命中 其他推理 :reasoning(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.07298 2019-03-26 cs.AI 74%

Preference Reasoning in Matching Procedures: Application to the Admission Post-Baccalaureat Platform

Youssef Hamadi, Souhila Kaci

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.07255 2017-10-02 cs.AI 74%

Assumption-Based Approaches to Reasoning with Priorities

Jesse Heyninck, Christian Straßer, Pere Pardo

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments Forthcoming in the proceedings of AI^3

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/9809013 2009-11-30 cs.AI cs.LO 74%

Reasoning about Noisy Sensors and Effectors in the Situation Calculus

Fahiem Bacchus, Joseph Y. Halpern, Hector J. Levesque

专题命中 其他推理 :reasoning(title);分类 cs.AI

Comments A preliminary version of the paper appeared in IJCAI '95

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08135 2025-11-12 cs.DC cs.AR 73%

UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing

Zhuoheng Ran, Chong Wu, Renjie Xu, Maolin Che, Hong Yan

专题命中 其他推理 :reasoning(title,comments)

Comments Accepted on 24 September 2025 at NeurIPS 2025 Efficient Reasoning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15964 2026-08-18 cs.CL cs.AI 新提交 73%

LLMs Get Smarter from Targeted Synthetic Multilingual Data

大语言模型(LLMs)从针对性的合成多语言数据中提升性能

Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke

机构 * UIUC(伊利诺伊大学厄巴纳-香槟分校) Uniphore(优尼佛(Uniphore)公司)

专题命中 其他推理 :reasoning(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本研究提出HOTFIXR数据生成框架,通过探测学生模型多语言弱点生成合成多语言数据,可提升大语言模型的多语言性能,降低微调引发的灾难性遗忘。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14390 2026-08-14 cs.CL cs.AI 版本更新 73%

REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models

REHEARSE:大语言模型中用于语言置信度校准的经验式排练

Ke Fang, Tianyi Zhao, Qianwen Wang, Lu Cheng

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Southern California(南加州大学) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 针对大语言模型置信度与实际正确性不匹配的问题,提出无训练的Rehearse方法,通过置信度校准博弈的反馈生成校准信号,在多模型多基准实验中显著降低了预期校准误差并提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09972 2026-08-12 physics.ao-ph cs.AI cs.LG 新提交 73%

Do AI weather models miss extremes?

AI气象模型是否遗漏极端天气?

Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey, Niels Poulsen, Philipp Seitz, Olivier Lam

专题命中 其他推理 :reasoning(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本研究验证11种物理与AI气象模型,发现极端天气下相对技能缺失是特定模型的特性而非AI气象模型整体的问题,部分AI模型在极端场景表现优于数值模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28335 2026-07-28 cs.CY cs.AI cs.CL 版本更新 73%

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

LLM-意识形态可塑性:将LLM的政治行为测量为上下文条件分布

Adib Sakhawat, Syed Rifat Raiyan, Tahsin Islam, Takia Farhin, Hasan Mahmud, Md Kamrul Hasan

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 通过系统实证,证明LLM的政治意识形态是上下文条件分布而非固定点,使用VAA-CHES投影模型在六个上下文轴上评估九个LLM,发现其对上下文高度敏感,但整体占据狭窄的Overton窗口。

Comments Under review, 40 pages, 18 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03333 2026-07-07 cs.DC cs.AI cs.LG 新提交 73%

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

SPORK:自我推测性分叉以加速智能体大语言模型推理

Huajun Bai, Weiwei Lv, Huichuan Zheng, Youyou Lu, Jiwu Shu

机构 * Tsinghua University(清华大学) Meituan(美团)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 研究LLM智能体推理中串行循环耗时问题,提出SPORK方法,利用模型自身预测提前调度推测工具调用,通过成本模型等组件优化,大幅提升推理效率且不影响任务准确性。

Comments 16 pages, 15 figures. Code: https://github.com/baihuajun24/spork

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00415 2026-07-02 cs.CL cs.LG 新提交 73%

A Mechanistic View of Authority Hierarchy in LLM Sycophancy

LLM 谄媚中权威层级的机制性视角

Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen

机构 * Independent Research(独立研究)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 通过受控医疗QA实验,发现LLM按感知权威程度分级响应,机制是特定后期层中正确答案表征被权威信号主动擦除,且该擦除与权威水平成比例、抵抗均值向量干预、仅部分可逆。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26071 2026-06-29 cs.LG cs.AI 新提交 73%

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

模型取证:调查令人担忧的行为是否反映对齐失败

Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan, Neel Nanda

机构 * MATS

专题命中 其他推理 :CoT(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出一个两步基线协议,通过读取思维链生成假设并编辑环境测试假设,以区分模型行为是出于恶意意图还是良性原因,并在六个环境中验证了该方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20382 2026-06-23 cs.CL cs.AI 版本更新 73%

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

言行不一:大型语言模型中的指令诱导冲突

Carolina Camassa, Derek Shiller

机构 * Future Impact Group(未来影响组)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了大型语言模型在面对指令与模式完成之间的冲突时的表现,发现指令遵循率在不同模型和指令下差异显著,输出多样性是预测鲁棒性的主要因素。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05150 2026-06-15 cs.CL cs.AI 版本更新 73%

Chronological Thinking in Full-Duplex Spoken Dialogue Language Models

全双工口语对话语言模型中的时间顺序思考

Donghang Wu, Haoyang Zhang, Chen Chen, Tianyu Zhang, Fei Tian, Xuerui Yang, Gang Yu, Hexin Liu, Nana Hou, Yuchen Hu, Eng Siong Chng

机构 * Nanyang Technological University(南洋理工大学) StepFun Mila

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 提出Chronological Thinking机制,让全双工对话模型在听用户说话时增量推理,不增加延迟,提升响应质量。

Comments Accepted by SIGDIAL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08446 2026-06-09 cs.LG cs.AI 新提交 73%

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models

Sparrow: 用于大语言模型稳定高效长上下文强化学习的稀疏 rollout

Yang Zhou, Ranajoy Sadhukhan, Zhaofeng Sun, Zhuoming Chen, Souvik Kundu, Saket Dingliwal, Sai Muralidhar Jayanthi, Aram Galstyan, Haizhong Zheng, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) Cornell University(康奈尔大学) Intel(英特尔) Amazon AGI(亚马逊AGI)

专题命中 其他推理 :CoT(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 针对RLVR中长上下文rollout计算昂贵的问题,提出Sparrow方法,通过动态稀疏度调度保持token级策略失配的下尾统计量稳定,在Qwen3系列模型上实现2.0-2.4倍加速,并推广到更大模型和编程领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30686 2026-06-01 cs.CR cs.AI cs.LG 73%

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

工具调用ReAct代理中深度相关的间接提示注入:注入深度、载荷框架和轮次预算敏感性

Mohammadreza Rashidi

机构 * Department of Computer Science(计算机科学系) AI and Media Analysis Lab(人工智能与媒体分析实验室) Berlin, Germany(柏林,德国)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 通过四个对照实验(共460次试验),研究在工具调用ReAct代理中,注入深度、载荷框架和轮次预算对间接提示注入攻击成功率的影响,发现注入深度是主导变量,且仅清理第一个工具观察可捕获67%的注入成功。

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16590 2026-05-22 cs.LG cs.AI q-bio.BM 73%

Atom-anchored LLMs speak Chemistry: A Retrosynthesis Demonstration

原子锚定的大语言模型:化学 retrosynthesis 的演示

Alan Kai Hassen, Andrius Bernatavicius, Antonius P. A. Janssen, Mike Preuss, Gerard J. P. van Westen, Djork-Arné Clevert

机构 * Machine Learning Research(机器学习研究) Pfizer Research and Development(辉瑞研发) Leiden Institute of Advanced Computer Science(莱顿高级计算机科学研究所) Leiden University(莱顿大学) Leiden Academic Centre for Drug Research(莱顿药物研究中心) Leiden Institute of Chemistry(莱顿化学研究所)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出了一种利用通用大语言模型进行分子推理的框架,通过原子标识符将链式推理与分子结构锚定,无需任务特定的模型训练,在单步 retrosynthesis 任务中实现了高成功率。

Comments Alan Kai Hassen and Andrius Bernatavicius contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18820 2026-05-20 cs.LG cs.AI 73%

Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision

前沿叠加的涌现:莫比乌斯吸引子与级联监督

Hongyu Gu, Jingwen Fu

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了通过叠加实现深度推理的问题,提出莫比乌斯吸引子和级联监督方法,证明了在Erdős-Rényi图上,叠加推理的涌现是通过建筑和监督的贡献实现的。

Comments 40 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03295 2026-05-14 cs.CL cs.AI cs.CY 73%

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

语言模型的目标选择与人类在自我指导学习任务中不同

Gaia Molinaro, Dave August, Danielle Perszyk, Anne G. E. Collins

机构 * University of California, Berkeley(加州大学伯克利分校) Amazon AGI Lab(亚马逊人工智能实验室)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 研究通过控制实验发现语言模型在目标选择上与人类存在显著差异,多数模型表现不佳,而人类则能逐步探索并达成多样化目标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08060 2026-05-11 cs.CL cs.AI cs.GT cs.MA 73%

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

记忆诅咒:扩展回忆如何在LLM代理中侵蚀合作意图

Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, Vincent Conitzer

机构 * Carnegie Mellon University(卡内基梅隆大学) Foundations of Cooperative AI Lab (FOCAL)(合作人工智能基础实验室) University of Michigan(密歇根大学) Harvard University(哈佛大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 研究发现扩展回忆会系统性地削弱多代理社会困境中的合作意图,通过三种分析揭示记忆内容对合作的影响,证明记忆是影响多代理行为的主动因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16817 2026-04-28 cs.LG cs.AI 73%

Self-Reinforcing Controllable Synthesis of Rare Relational Data via Bayesian Calibration

通过贝叶斯校准实现自我强化的可控稀有关系数据合成

Chongsheng Zhang, Hao Wang, Zelong Yu, Esteban Garces Arias, Julian Rodemann, Zhanshuo Zhang, Qilong Li, Gaojuan Fan, Krikamol Muandet, Christian Heumann

机构 * Henan University, China(河南大学) Department of Statistics, LMU Munich(慕尼黑大学统计系) MCML Munich(慕尼黑MCML) CISPA Helmholtz Center for Information Security, Saarbrücken, Germany(萨尔布吕肯德国海德堡中心信息安全研究所)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

AI总结 本文提出RDDG框架,通过动态引导生成表格数据以提升不平衡分类性能,采用核心集选择和自强化反馈机制优化生成质量。

Comments Accepted at: Findings of the Association for Computational Linguistics: ACL 2026 (ACL 2026 Findings), San Diego, California, USA, July 2-7, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10539 2026-04-14 cs.LG cs.AI 73%

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs

IceCache: 用于长序列LLM的内存高效KV缓存管理

Yuzhen Mao, Qitong Wang, Martin Ester, Ke Li

机构 * Simon Fraser University(西蒙菲莎大学) Harvard University(哈佛大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本文提出IceCache,通过语义令牌聚类与PagedAttention结合,提升长序列LLM的内存效率与性能,实验表明其在256个令牌预算下保持99%的准确性,且在延迟和精度上优于其他方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08425 2026-04-10 cs.AI cs.CL 73%

Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM

学习谁不一致:用于建模标注者分布的群体重要性加权

Samay U. Shetty, Tharindu Cyril Weerasooriya, Deepak Pandita, Christopher M. Homan

机构 * Rochester Institute of Technology(罗切斯特理工学院)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 DiADEM通过学习群体重要性权重,有效建模标注者分歧,优于现有方法,在多个基准上取得显著成果,揭示种族和年龄对标注者分歧的关键影响。

详情

展开后加载摘要…

URL PDF HTML 收藏