arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5768 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5768 篇

2407.00071 2024-07-02 cs.AI cs.CL cs.ET cs.LG 85%

Combinatorial Reasoning: Selecting Reasons in Generative AI Pipelines via Combinatorial Optimization

Mert Esencan, Tarun Advaith Kumar, Ata Akbari Asanjan, P. Aaron Lott, Masoud Mohseni, Can Unlu, Davide Venturelli, Alan Ho

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03204 2024-05-21 eess.AS cs.AI cs.CL cs.LG cs.SD 85%

RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Detai Xin, Xu Tan, Kai Shen, Zeqian Ju, Dongchao Yang, Yuancheng Wang, Shinnosuke Takamichi, Hiroshi Saruwatari, Shujie Liu, Jinyu Li, Sheng Zhao

专题命中 其他推理 :chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16258 2025-08-26 cs.CL cs.AI cs.CV 84%

IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection

Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee

机构 * College of Information Sciences and Technology(信息科学与技术学院) The Pennsylvania State University(宾夕法尼亚州立大学) Department of Computer Science and Engineering(计算机科学与工程系) National Institute of Technology, Tiruchirappalli(特里奇里帕利理工学院)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments Accepted in the COLM First Workshop on Pragmatic Reasoning in Language Models (PragLM), Montreal, Canada, October 2025, https://sites.google.com/berkeley.edu/praglm

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08505 2024-04-03 cs.CL 84%

Semi-Structured Chain-of-Thought: Integrating Multiple Sources of Knowledge for Improved Language Model Reasoning

Xin Su, Tiep Le, Steven Bethard, Phillip Howard

专题命中 其他推理 :reasoning(title);chain-of-thought(title);分类 cs.CL

Comments NAACL 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09867 2026-08-11 cs.CR cs.AI cs.LG 新提交 84%

Stealing Reasoning Traces from Proprietary LLM APIs

从专有大语言模型API窃取推理轨迹

Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 该研究发现专有LLM API的加密推理轨迹存在跨会话、用户及模型的兼容性漏洞,开发了可扩展解密越狱方法,能提取模型推理、窃取PII等,还提出了对应缓解措施。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29287 2026-08-03 cs.CL cs.AI 新提交 84%

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

带思考的翻译:面向多领域机器翻译的难度自适应推理强化学习方法

Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen, Xiaodong Shi

机构 * Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院) School of Informatics, Xiamen University(厦门大学信息学院) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism(文化和旅游部闽台非物质文化遗产数字化保护与智能处理重点实验室(厦门大学))

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对多领域机器翻译的难度差异挑战,提出TwT框架,经两阶段训练后,其7B和14B参数版本在翻译质量上优于更大的SOTA模型,且token使用量降低32%-60%。

Comments 34 pages, 17 figures, and 21 tables. Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19825 2026-07-28 cs.CL cs.AI 84%

Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost

简洁思考:输出长度对LLM推理和成本的影响

Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, Fabrizio Giacomelli

专题命中 其他推理 :reasoning(title);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CCoT方法,通过控制输出长度提升LLM推理的简洁性和效率,并验证了相关指标的有效性。

Comments Preprint version, under review

Journal ref Information Sciences 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13903 2026-07-16 cs.MA cs.AI cs.LG 版本更新 84%

Benefits and Limitations of Communication in Multi-Agent Reasoning

多智能体推理中通信的益处与局限性

Michael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi, Guillaume Rabusseau, Michael Hahn

机构 * Mila & Université de Montréal(Mila与蒙特利尔大学) University of Oxford(牛津大学) Stanford University(斯坦福大学) Saarland University(萨尔兰大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 研究多智能体推理中通信的益处与局限,提出理论框架分析其表达能力,应用于三个算法族,得出相关界限,通过实验验证关键量权衡,为设计可扩展多智能体推理系统提供指导。

Comments 34 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10184 2026-06-10 cs.LG cs.AI 新提交 84%

Dropout-GRPO: Variational Stochasticity for Continuous Latent Reasoning

Dropout-GRPO: 用于连续潜在推理的变分随机性

Wooil Jung

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 针对GRPO在连续潜在推理模型中因确定性轨迹导致优势为零的问题,提出通过结构化Dropout引入随机性,使GRPO能优化贝叶斯模型平均策略,在GSM8K上提升Coconut基线准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05988 2026-06-05 cs.LG cs.CL 84%

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

压缩-蒸馏:面向高效知识蒸馏的推理轨迹压缩

Maxime Griot, Paul Steven Scotti, Tanishq Mathew Abraham

机构 * Université catholique de Louvain(列日天主教大学) Sophont Inc(Sophont公司)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 本文提出在知识蒸馏前对推理轨迹进行事后压缩,以降低训练成本并缩短推理输出,实验表明压缩在准确率与效率间存在权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05025 2026-06-04 cs.LG cs.AI 84%

Invariant Gradient Alignment for Robust Reasoning Distillation

不变梯度对齐用于鲁棒推理蒸馏

Zehua Cheng, Wei Dai, Jiahao Sun

机构 * University of Oxford(牛津大学) FLock.io

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 提出不变梯度对齐(IGA)框架,通过逻辑同构集、连续梯度冲突掩码和截断SVD投影,对齐不同语义域但逻辑结构相同的梯度更新,提升大语言模型在分布外输入上的鲁棒性。

Comments 30 Pages

Journal ref In Proceedings of European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01042 2026-06-02 cs.LG cs.AI 84%

Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning

似真性不是预测:基于LLM的细胞扰动推理的对比证据

Xinyu Yuan, Xixian Liu, Jianan Zhao, Yashi Zhang, Hongyu Guo, Jian Tang

机构 * Mila - Québec AI Institute(魁北克人工智能研究所) University of Montréal(蒙特利尔大学) HEC Montréal(蒙特利尔HEC商学院) University of Ottawa(渥太华大学) National Research Council of Canada(加拿大国家研究理事会) CIFAR AI Chair(CIFAR人工智能 chair)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文发现基于大语言模型的细胞扰动推理虽能生成生物上合理的解释,但实际预测性能差,并提出CORE方法通过对比证据组织来提升扰动特异性预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21357 2026-04-24 cs.AI cs.CL 84%

ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs

ReaGeo:基于大语言模型的推理增强端到端地名编码

Jian Cui, Zhiyuan Ren, Desheng Weng, Yongqi Zhao, Gong Wenbin, Yu Lei, Zhenning Dong

机构 * Amap, Alibaba Group(阿里巴巴集团阿地图) Tsinghua University(清华大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 ReaGeo利用大语言模型解决传统多阶段方法在地理数据库中依赖文本或向量相似性检索的局限性,通过将坐标转换为geohash序列,引入链式推理机制提升空间关系推理能力,并通过距离偏差奖励的强化学习优化生成精度。

Comments 12 pages, 8 figures, submitted to ACM SIGSPATIAL 2024 (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07663 2026-04-22 cs.AI cs.CL 84%

Reasoning Models Will Sometimes Lie About Their Reasoning

推理模型有时会谎称其推理过程

William Walden, Miriam Wanner

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 其他推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了在模型被明确提示可能接收异常输入的情况下,推理模型的忠实性。发现尽管提示能提升原有指标,但新提出的更细致指标显示模型常否认使用提示,挑战了推理监控和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25960 2026-03-30 cs.CL cs.AI 84%

When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models

当链式思维产生反效果:评估医疗语言模型在提示敏感性上的表现

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven(纽黑文大学SAIL实验室)

专题命中 其他推理 :chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 研究评估医疗语言模型在不同提示格式下的鲁棒性,发现链式思维提示降低准确性,少样本示例和答案选项洗牌显著影响性能,且领域特定模型的提示工程策略不适用于通用模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18453 2026-03-03 cs.AI cs.CL 84%

Reason Like a Radiologist: Chain-of-Thought and Reinforcement Learning for Verifiable Report Generation

像放射科医生一样推理:用于可验证报告生成的推理链和强化学习

Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Huichi Zhou, Zhengqing Yuan, Zhifan Gao, Lei Zhu, Giorgos Papanastasiou, Yingying Fang, Guang Yang

专题命中 其他推理 :chain-of-thought(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 BoxMed-RL通过推理链和强化学习生成可验证的放射科报告,提升报告的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25666 2026-02-11 cs.LG cs.CL 84%

Nudging the Boundaries of LLM Reasoning

推动大语言模型推理边界的 nudging

Justin Chih-Yao Chen, Becky Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang, Jiaxin Zhang, Mohit Bansal, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究)

专题命中 其他推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.LG

AI总结 NuRL通过自动生成提示提升大语言模型推理上限,解决传统RL无法学习无法解决样本的问题。

Comments ICLR 2026 (Camera-Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00173 2026-02-03 cs.LG cs.AI 84%

Learning Robust Reasoning through Guided Adversarial Self-Play

通过引导对抗自博弈学习鲁棒推理

Shuozhe Li, Vaishnav Tadiparthi, Kwonjoon Lee, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Lizhang Chen, Amy Zhang, Liu Leqi

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Honda Research Institute USA(本田美国研究院)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 GASP通过引导对抗自博弈方法提升推理模型的鲁棒性,使其在面对误导和扰动上下文时保持稳定并提高准确度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11517 2026-01-19 cs.CL cs.AI 84%

Do explanations generalize across large reasoning models?

解释是否能在大型推理模型中泛化?

Koyena Pal, David Bau, Chandan Singh

专题命中 其他推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了大型推理模型生成的解释是否能泛化,发现CoT解释能提高模型间一致性,并提出基于句子的融合策略提升一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05560 2026-01-12 cs.CL cs.AI 84%

ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging

ReasonAny: 通过简单有效的模型融合将推理能力融入任意模型

Junyao Yang, Chen Qian, Dongrui Liu, Wen Shen, Yong Liu, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) Renmin University of China(中国人民大学) Tongji University(同济大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 ReasonAny通过对比梯度识别解决推理与领域性能崩溃问题,有效融合推理能力与领域专长,提升多领域模型性能。

Comments 22 pages, 6 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03542 2026-01-08 cs.CL cs.AI 84%

Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models

层序倒置:重新思考大语言模型中的潜在多跳推理

Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, Qi Liu

机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 本文提出"层序倒置"现象,通过概率性回忆与提取框架解释大语言模型中多跳推理的机制,揭示了层序倒置与总跳步数的关系,并提供了多跳失败的诊断方法。

Comments 16 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15219 2025-12-18 cs.CL cs.AI 84%

RFKG-CoT: Relation-Driven Adaptive Hop-count Selection and Few-Shot Path Guidance for Knowledge-Aware QA

RFKG-CoT: 基于关系的自适应步数选择与少样本路径引导用于知识感知问答

Chao Zhang, Minghan Li, Tianrui Lv, Guodong Zhou

机构 * Chao Zhang, Minghan Li, Tianrui Lv, Guodong Zhou(张超,李明翰,吕天睿,周国栋)

专题命中 其他推理 :CoT(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 RFKG-CoT通过关系驱动的自适应步数选择和少样本路径引导,提升知识感知问答的准确性与可靠性。

Comments 9pages, 5 figures, accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13157 2025-10-16 cs.CE cs.AI cs.CL 84%

Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval

Subhendu Khatuya, Shashwat Naidu, Pawan Goyal, Niloy Ganguly

机构 * Indian Institute of Technology Kharagpur(印度理工学院卡格鲁分校)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments This work has been accepted for publication in the Main Conference of the Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15969 2025-10-16 cs.LG cs.CL 84%

LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning

Haoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang, Song Guo

机构 * HKUST(香港科技大学) HK PolyU(香港理工大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 84%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 其他推理 :reasoning(title,abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17387 2025-08-29 cs.LG cs.AI 84%

Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoning

Yicong Wu, Guangyue Lu, Yuan Zuo, Huarong Zhang, Junjie Wu

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04374 2025-06-06 cs.AI cs.CL 84%

A Statistical Physics of Language Model Reasoning

Jack David Carson, Amir Reisizadeh

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20129 2025-06-04 cs.CL cs.LG 84%

Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking

Yifan Zhang, Wenyu Du, Dongming Jin, Jie Fu, Zhi Jin

机构 * Key Laboratory of High Confidence Software Technology (PKU), MOE, China(高性能软件技术关键实验室(PKU),教育部,中国) School of Computer Science, Peking University, China(北京大学计算机学院,中国) The University of Hong Kong(香港大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 其他推理 :chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18283 2025-05-27 cs.CL cs.AI cs.MA 84%

TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification

Jianghao Wu, Feilong Tang, Yulong Li, Ming Hu, Haochen Xue, Shoaib Jameel, Yutong Xie, Imran Razzak

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) Monash University(墨尔本大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Southampton(南安普顿大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments 16 pages including references, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14038 2025-05-21 cs.AI cs.CL 84%

ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data

Xinzhe Zheng, Sijie Ji, Jiawei Sun, Renqi Chen, Wei Gao, Mani Srivastava

机构 * California Institute of Technology(加州理工学院) UCLA(加州大学洛杉矶分校) National University of Singapore(新加坡国立大学) Hangzhou Dianzi University(杭州电子科技大学) Fudan University(复旦大学)

专题命中 其他推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏