arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5792 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5792 篇

2601.10274 2026-01-16 cs.LG cs.AI cs.IT cs.NI math.IT math.OC 76%

Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers

面向队列的推理令牌优化:在LLM服务器中准确率与延迟的权衡

Emre Ozbas, Melih Bastopcu

机构 * Department of Electrical and Electronics Engineering(电气与电子工程系)

专题命中 其他推理 :reasoning(title);分类 cs.AI、cs.LG

AI总结 本文提出了一种面向队列的推理令牌优化方法,通过优化LLM服务器中准确率与延迟的权衡,提升系统性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12701 2025-12-16 cs.CV cs.CL cs.LG 76%

Efficient Vision-Language Reasoning via Adaptive Token Pruning

通过自适应令牌修剪实现高效的视觉语言推理

Xue Li, Xiaonan Song, Henry Hu

机构 * Scholar42(学者42) InfiniPouch LLC(InfiniPouch公司) Labelbox, Inc.(Labelbox公司)

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.LG

AI总结 本研究提出自适应令牌修剪方法,通过动态保留关键令牌提升视觉语言模型的推理效率和稳定性,减少计算量并保持精度。

Comments 10 pages, 3 figures. Expanded version of an extended abstract accepted at NeurIPS 2025 Workshop on VLM4RWD. Presents methodology and preliminary experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10004 2025-09-15 cs.CL cs.AI 76%

Unsupervised Hallucination Detection by Inspecting Reasoning Processes

Ponhvoan Srey, Xiaobao Wu, Anh Tuan Luu

机构 * Nanyang Technological University(南洋理工大学)

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

Comments To appear in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08235 2025-07-14 cs.LG cs.AI 76%

InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems

Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Rajiv Ramnath

机构 * Department of Computer Science(计算机科学系) Engineering, Ohio State University, Ohio, US(工程学院,俄亥俄州立大学,俄亥俄州,美国)

专题命中 其他推理 :reasoning(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03327 2025-07-08 cs.CL cs.AI 76%

Read Quietly, Think Aloud: Decoupling Comprehension and Reasoning in LLMs

Yuanxin Wang, Ganesh Venkatesh

机构 * AppliedML

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10886 2025-03-17 cs.CV cs.AI cs.IR cs.LG q-bio.PE 76%

Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification

Nathaniel Lesperance, Sujeevan Ratnasingham, Graham W. Taylor

专题命中 其他推理 :reasoning(title);分类 cs.AI、cs.LG

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08460 2025-01-16 cs.CV cs.AI cs.CL 76%

Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time

Mihai Masala, Marius Leordeanu

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11009 2024-12-17 cs.AI cs.CL cs.CY 76%

Dual Traits in Probabilistic Reasoning of Large Language Models

Shenxiong Li, Huaxia Rui

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14744 2024-10-22 cs.CL cs.AI 76%

Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors

Anthony Sicilia, Malihe Alikhani

专题命中 其他推理 :chain-of-thought(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04562 2024-04-11 cs.CL cs.AI 76%

Towards Foundation Models for Knowledge Graph Reasoning

Mikhail Galkin, Xinyu Yuan, Hesham Mostafa, Jian Tang, Zhaocheng Zhu

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15631 2024-02-27 cs.CL cs.AI 76%

Fine-Grained Self-Endorsement Improves Factuality and Reasoning

Ante Wang, Linfeng Song, Baolin Peng, Ye Tian, Lifeng Jin, Haitao Mi, Jinsong Su, Dong Yu

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14435 2023-10-24 cs.CL cs.AI 76%

Retrieval-Augmented Chain-of-Thought in Semi-structured Domains

Vaibhav Mavi, Abulhair Saparov, Chen Zhao

专题命中 其他推理 :chain-of-thought(title);分类 cs.CL、cs.AI

Comments to appear in NLLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03834 2021-10-07 cs.AI cs.LG 76%

Bob and Alice Go to a Bar: Reasoning About Future With Probabilistic Programs

David Tolpin, Tomer Dobkin

专题命中 其他推理 :reasoning(title);分类 cs.AI、cs.LG

Comments 31 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.05922 2018-04-18 cs.CL cs.LG 76%

Neural Models for Reasoning over Multiple Mentions using Coreference

Bhuwan Dhingra, Qiao Jin, Zhilin Yang, William W. Cohen, Ruslan Salakhutdinov

专题命中 其他推理 :reasoning(title);分类 cs.CL、cs.LG

Comments NAACL 2018 (Short Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11521 2025-10-28 cs.CR cs.AI 76%

Detecting Various DeFi Price Manipulations with LLM Reasoning

Juantao Zhong, Daoyuan Wu, Ye Liu, Maoyi Xie, Yang Liu, Yi Li, Ning Liu

机构 * Daoyuan Wu is the co-first author.(达元吴是共同第一作者) Ye Liu is the corresponding author.(叶刘是通讯作者)

专题命中 其他推理 :reasoning(title,comments);分类 cs.AI

Comments Accepted by ASE 2025. Please cite the conference version of this paper, e.g., "Juantao Zhong, Daoyuan Wu, Ye Liu, Maoyi Xie, Yang Liu, Yi Li, Ning Liu. Detecting Various DeFi Price Manipulations with LLM Reasoning. In 40th IEEE/ACM International Conference on Automated Software Engineering (ASE 2025)"

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18930 2024-06-28 cs.AI cs.DM cs.LO cs.SC 76%

Reasoning About Action and Change

Florence Dupin de Saint-Cyr, Andreas Herzig, Jérôme Lang, Pierre Marquis

专题命中 其他推理 :reasoning(title,journal_ref);分类 cs.AI

Journal ref Marquis, Pierre; Papini, Odile; Prade, Henri. A Guided Tour of Artificial Intelligence Research, 1 / 3, Springer International Publishing, pp.487-518, 2020, Knowledge Representation, Reasoning and Learning, 978-3-030-06163-0

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30649 2026-07-01 cs.CY 新提交 75%

Thinking Out Loud: Real-Time Deception Monitoring in Asymmetric LLM Negotiations

大声思考:非对称LLM谈判中的实时欺骗监控

Nolan Coffey, Faithful Odoi, Makenzie Johnson, Nasir U. Eisty

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 提出轻量级实时思维链监控器,在二手车销售场景中检测卖家隐藏缺陷的欺骗行为,实验表明监控能提高买家退出率但存在智能差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05386 2026-06-30 cs.LG cs.AI cs.CL 75%

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training

强化微调自然缓解持续训练中的遗忘

Song Lai, Haohan Zhao, Rong Feng, Changyi Ma, Wenzhuo Liu, Hongbo Zhao, Xi Lin, Dong Yi, Qingfu Zhang, Hongbin Liu, Gaofeng Meng, Fei Zhu

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文比较了监督微调与强化微调在持续训练中的影响,发现强化微调能有效保留先验知识,优于监督微调和多任务训练,且在标准基准上提升模型通用知识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10352 2026-06-03 cs.CL cs.AI cs.LG 75%

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

从可解释性工件中学习自我解释:在向量-标签对上训练轻量级适配器

Keenan Pepper, Alex McKenzie, Florin Pop, Stijn Servaes, Martin Leitgab, Mike Vaiana, Judd Rosenblatt, Michael S. A. Graziano, Diogo de Lucena

机构 * University of Washington(华盛顿大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过训练轻量级适配器(标量仿射适配器,仅需d_model+1参数)在可解释性工件上,保持语言模型完全冻结,实现了跨任务和模型族的可靠自我解释,在稀疏自编码器特征标注、主题识别和多跳推理桥接实体解码等任务上显著优于未训练基线。

Comments 26 pages, 18 tables, 17 figures. Code and data at https://github.com/agencyenterprise/selfie-adapters

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29601 2026-05-29 cs.CL cs.AI cs.LG 75%

Training Deliberative Monitors for Black-Box Scheming Detection

训练审慎监控器用于黑箱策划检测

Aditya Sinha, Akshat Naik, Victor Gillioz, Simon Storf, Kilian Merkelbach, Rich Barton-Cooper, Axel Højmark, Marius Hobbhahn

机构 * Independent(独立) MATS Research(MATS研究) Astra Fellowship Apollo Research(Apollo研究)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出一种基于行动轨迹的审慎监控方法,通过蒸馏前沿模型的推理过程训练开源模型,以低成本高精度检测智能体的策划与破坏行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10701 2026-04-14 cs.LG cs.AI cs.CL 75%

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning

带回价值模型:生成式批评者在大语言模型强化学习中的价值建模

Zikang Shan, Han Zhong, Liwei Wang, Li Zhao

机构 * Peking University(北京大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出生成式批评者GenAC,通过链式推理改进价值建模,提升RL性能。

Comments 16 pages including appendix, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06613 2026-04-10 cs.CL cs.AI cs.IT cs.LG math.IT 75%

The Detection-Extraction Gap: Models Know the Answer Before They Can Say It

检测-提取间隙:模型在能说出答案前就已经知道答案

Hanyang Wang, Mingxuan Zhu

机构 * The University of Chicago(芝加哥大学) Imperial College London(伦敦帝国学院)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现模型在确定答案后仍生成大量内容,揭示了检测与提取之间的结构性差距,并提出BAEE方法通过利用自由延续提高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04451 2026-02-06 cs.IR 75%

SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image Retrieval

SDR-CIR:一种基于链式推理的语义去偏检索框架用于无训练零样本复合图像检索

Yi Sun, Jinyu Xu, Qing Xie, Jiachen Li, Yanchun Ma, Yongjian Liu

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 SDR-CIR通过语义去偏排名方法提升无训练零样本复合图像检索性能。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06299 2026-01-23 cs.CY cs.AI cs.CL cs.LG 75%

How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare

恶意AI群如何威胁民主:代理AI与大语言模型的融合标志着信息战争的新前沿

Daniel Thilo Schroeder, Meeyoung Cha, Andrea Baronchelli, Nick Bostrom, Nicholas A. Christakis, David Garcia, Amit Goldenberg, Yara Kyrychenko, Kevin Leyton-Brown, Nina Lutz, Gary Marcus, Filippo Menczer, Gordon Pennycook, David G. Rand, Maria Ressa, Frank Schweitzer, Dawn Song, Christopher Summerfield, Audrey Tang, Jay J. Van Bavel, Sander van der Linden, Jonas R. Kunst

机构 * Department of Sustainable Communication Technologies, SINTEF Digital(可持续通信技术系,SINTEF数字) Max Planck Institute for Security and Privacy(安全与隐私研究所) Department of Mathematics, City St George’s University of London(数学系,圣乔治大学) Macrostrategy Research Initiative(战略研究计划) Human Nature Lab, Yale University(人性实验室,耶鲁大学) Department of Politics and Public Administration, University of Konstanz(政治与公共管理系,康斯坦茨大学) Harvard Business School, Harvard University(哈佛商学院,哈佛大学) Department of Psychology, University of Cambridge(心理学系,剑桥大学) Department of Computer Science, University of British Columbia(计算机科学系,不列颠哥伦比亚大学) Department of Human Centered Design & Engineering, University of Washington(以人为本设计与工程系,华盛顿大学) Department of Psychology, New York University(心理学系,纽约大学) Observatory on Social Media and Luddy School of Informatics, Computing, and Engineering, Indiana University(社交媒体观察所和信息、计算与工程学院,印第安纳大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文探讨了恶意AI群通过融合代理AI与大语言模型对民主构成的威胁,并提出多方面的干预措施。

Comments 5 Pages, This is the author's version of the work. It is posted here by permission of the AAAS for personal use, not for redistribution. The definitive version was published in Science on January 22, 2026, DOI: 10.1126/science.adz1697

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23002 2025-12-05 cs.CV 75%

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

JarvisEvo: 向自进化式照片编辑代理迈进:协同编辑器-评估器优化

Yunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin, Kaixiong Gong, Wenbo Li, Bin Lin, Zhenxi Li, Shiyi Zhang, Yuyang Peng, Wenxun Dai, Xinghao Ding, Chunyu Wang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯文元)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 JarvisEvo通过协同编辑器-评估器优化,实现自进化式照片编辑代理,提升编辑质量和内容保真度。

Comments 31 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02025 2025-11-26 cs.SE 75%

SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash Data

SAFE: 借助LLM从多模态碰撞数据中驱动场景的ADS测试

Siwei Luo, Yang Zhang, Yao Deng, Linfeng Liang, Xi Zheng

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

AI总结 SAFE通过多模态提取和LLM技术,提升从碰撞数据中重建真实场景和检测安全违规的能力,实现高准确率和高效生成。

Comments The paper has been accepted for publication in the proceedings of the IEEE/ACM 48th International Conference on Software Engineering to be held 12-18 April 2026 (ICSE2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25195 2025-10-30 cs.SE 75%

Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models

Shuochuan Li, Zan Wang, Xiaoning Du, Zhuo Wu, Jiuqiao Yu, Junjie Chen

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13500 2025-10-14 cs.CL cs.AI cs.LG 75%

Noise Injection Systemically Degrades Large Language Model Safety Guardrails

Prithviraj Singh Shahani, Kaveh Eskandari Miandoab, Matthias Scheutz

机构 * Tufts University(塔夫茨大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24663 2025-09-30 cs.CL cs.AI cs.LG 75%

InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation

Weilin Zhao, Zihan Zhou, Zhou Su, Chaojun Xiao, Yuxuan Li, Yanghao Li, Yudi Zhang, Weilun Zhao, Zhen Li, Yuxiang Huang, Ao Sun, Xu Han, Zhiyuan Liu

机构 * Tsinghua University(清华大学) OpenBMB(开放大脑实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13329 2025-09-24 cs.CL cs.AI cs.LG 75%

Language Models Can Predict Their Own Behavior

Dhananjay Ashok, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) University of Southern California(南加州大学)

专题命中 其他推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏