arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 5768 信号源:cs.CL, cs.AI, cs.LG

1. 其他推理 5768 篇

2605.15041 2026-05-15 cs.AI cs.CL 81%

Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use

基于案例的自适应推理与执行校准:大型语言模型工具使用

Renning Pang, Tian Lan, Leyuan Liu, Piao Tong, Sheng Cao, Xiaosong Zhang

机构 * University of Electronic Science and Technology of China(电子科技大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出CAST框架,通过历史执行轨迹作为结构化案例,提取复杂性和失败特征以优化推理策略,提升工具使用准确性并减少冗余推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13687 2026-05-14 cs.LG cs.AI stat.ML 81%

A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning

具有可预测扩展规律和推理证明益处的分层语言模型

Jason Gaitonde, Frederic Koehler, Elchanan Mossel, Joonhyung Shin, Allan Sly

机构 * Duke University(杜克大学) University of Chicago(芝加哥大学) Massachusetts Institute of Technology(麻省理工学院) Princeton University(普林斯顿大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种合成语言家族,通过树上的广播过程生成,分析上下文长度和推理在自回归生成中的作用。通过精确k-gram假设,证明了在特定条件下,上下文深度与序列生成的方差和峰度关系,展示了推理模型在有限内存下能精确生成真实语言。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12518 2026-05-14 cs.CL cs.AI 81%

TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models

TimelineReasoner: 通过大推理模型推进时间线摘要

Liancheng Zhang, Xiaoxi Li, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出TimelineReasoner框架,通过大推理模型的主动推理能力,改进时间线摘要任务,提升准确性、覆盖性和连贯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10781 2026-05-12 cs.LG cs.CL 81%

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

反叛学生:通过自蒸馏RLVR反转教师信号进行推理探索

Jeonghye Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang

机构 * Microsoft Research(微软研究院) KAIST(韩国科学技术院)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出RLRT,通过反转自蒸馏信号增强正确路径上的推理,提升RLVR性能,建立信息不对称作为新设计轴。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01569 2026-05-11 cs.AI cs.CL 81%

InvThink: Premortem Reasoning for Safer Language Models

InvThink:用于更安全语言模型的预mortem推理

Yubin Kim, Taehan Kim, Eugene Park, Chunjong Park, Cynthia Breazeal, Daniel McDuff, Hae Won Park

机构 * MIT(麻省理工学院) Google Research(谷歌研究院) Google DeepMind(谷歌DeepMind) Samsung Research(三星研究)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 InvThink通过预mortem推理框架提升语言模型安全性,通过枚举、分析和约束潜在故障来生成响应,相比现有方法在安全性和减少有害行为方面表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21137 2026-05-08 cs.CL cs.AI 81%

Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification

通过联合多任务学习增强科学课堂话语分析以进行推理组件分类

Jiho Noh, Mukhesh Raghava Katragadda, Raymond Carl, Soon Lee

机构 * Department of Computer Science, Kennesaw State University(肯纳邦大学计算机科学系) Bagwell College of Education, Kennesaw State University(肯纳邦大学巴格威尔教育学院)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出ADAS系统,通过联合多任务学习对教师和学生话语进行类型和推理组件分类,解决少数类标签不平衡问题,提升课堂话语分析效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24809 2026-04-29 cs.LG cs.AI 81%

Nautile-370M: Spectral Memory Meets Attention in a Small Reasoning Model

Nautile-370M:在小推理模型中融合频谱记忆与注意力

Maixent Chenebaux

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 Nautile-370M是一款3.7亿参数的小语言模型,通过融合频谱记忆与注意力机制,在有限参数和推理预算下实现高效推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15707 2026-04-29 cs.CL cs.AI 81%

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?

大型语言模型在推理任务上的表现是否受提问方式的影响?

Seok Hwan Song, Mohna Chakraborty, Qi Li, Wallapak Tavanapong

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨了不同提问方式对大型语言模型推理任务准确性的影响,发现问题类型显著影响模型表现,选项数量和用词选择也会影响最终答案选择的准确性。

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18936 2026-04-22 cs.LG cs.AI hep-ph hep-th 81%

Fine-Tuning Small Reasoning Models for Quantum Field Theory

对量子场论进行小规模推理模型的微调

Nathaniel S. Woodward, Zhiqi Gao, Yurii Kvasiuk, Kendrick M. Smith, Frederic Sala, Moritz Münchmeyer

机构 * Department of Physics, University of Wisconsin-Madison(威斯康星大学麦迪逊分校物理系) Department of Computer Science, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系) Perimeter Institute for Theoretical Physics(理论物理研究所)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在训练大型语言模型时,领域特定物理推理能力的发展,通过微调小规模推理模型,生成合成问题和人类编写的问题,进行强化学习和监督微调实验,分析推理错误的变化,并公开数据管道和训练数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15725 2026-04-20 cs.LG cs.AI 81%

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

通过语义触发和心理框架针对大推理模型的推理定向劫持攻击

Zehao Wang, Lanjun Wang

机构 * College of Intelligence and Computing(智能与计算学院) School of New Media and Communication(新媒体与传播学院) Shanghai Key Laboratory of Data Science(上海数据科学 key laboratory)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出PRJA框架,通过语义触发选择模块和心理指令生成模块,解决大推理模型推理过程中的安全问题,实验显示在五个问答数据集上攻击成功率达83.6%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14631 2026-04-17 cs.CL cs.AI 81%

StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation

StoryCoder: 为LLM代码生成中的结构化推理提供叙述改写

Geonhui Jang, Dongyoon Han, YoungJoon Yoo

机构 * Dept. of Artificial Intelligence, Chung-Ang University(Chung-Ang 大学人工智能系) NAVER AI Lab(NAVER AI 实验室)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 StoryCoder通过将代码生成问题转化为连贯的自然语言叙述,提升模型推理和规划能力,实验显示在多个数据集上平均提升18.7%的零样本准确率,且改进了算法策略和代码结构。

Comments 21 pages, 12 figures. ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14362 2026-04-17 cs.CL cs.AI cs.IR 81%

APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI

APEX-MEM: 基于时序推理的代理半结构化记忆用于长期对话AI

Pratyay Banerjee, Masud Moshtaghi, Shivashankar Subramanian, Amita Misra, Ankit Chadha

机构 * Amazon(亚马逊)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 APEX-MEM通过结合属性图、只追加存储和多工具检索代理,解决大语言模型在长期对话记忆中的可靠性问题,实现88.88%的LOCOMO问答准确率和86.2%的LongMemEval表现。

Comments Accepted to ACL 2026 Mains

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11791 2026-04-14 cs.LG cs.AI 81%

A Mechanistic Analysis of Looped Reasoning Language Models

循环推理语言模型的机理分析

Hugh Blayney, Álvaro Arroyo, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Michael M. Bronstein, Xiaowen Dong

机构 * University of Oxford(牛津大学) Mila – Quebec AI Institute(魁北克AI研究所)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文分析了循环推理语言模型中潜在状态的机理,揭示了循环块在潜在空间中的稳定轨迹及注意力头行为的稳定特性。

Comments 39 pages, 63 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10900 2026-04-14 cs.AI cs.LG 81%

CASK: Core-Aware Selective KV Compression for Reasoning Traces

CASK:面向推理轨迹的核心感知选择性KV压缩

Buseong Kim, Heejun Gwon

机构 * d’strict Korea(d’strict 韩国)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 CASK通过核心保护与选择性擦除策略优化推理轨迹的KV缓存,提升内存效率和推理稳定性,在AIME24和AIME25测试中优于TriAttention。

Comments 25 pages, 8 figures, 3 main tables, appendices included

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13909 2026-04-08 cs.CL cs.AI 81%

Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning

知识推理语言模型:统一知识与语言以进行归纳知识图谱推理

Xingrui Zhuo, Jiapu Wang, Gongqing Wu, Zhongyuan Wang, Jichen Zhang, Shirui Pan, Xindong Wu

机构 * The Key Laboratory of Knowledge Engineering with Big Data (the Ministry of Education of China), Hefei University of Technology, China(合肥工业大学大数据知识工程教育部重点实验室) School of Computer Science and Information Engineering, Hefei University of Technology, China(合肥工业大学计算机与信息学院) Nanjing University of Science and Technology, China(南京理工大学) China Unicom Digital Technology Co., Ltd., Beijing, China(联通数字科技有限公司) China Unicom Internet of Things Co., Ltd., Nanjing, China(联通物联网有限责任公司) Shandong Inspur Science Research Institute, Jinan, China(山东浪潮科学研究院) Griffith University, Australia(格里菲斯大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出KRLM,通过统一语言模型知识与知识图谱上下文,解决归纳知识图谱推理中的知识扭曲和生成幻觉问题,实验表明其在25个真实数据集上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00716 2026-04-02 cs.AI cs.LG 81%

CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection

CircuitProbe: 通过稳定性区检测预测Transformer中的推理电路

Rajkiran Panuganti

机构 * Independent Researcher(独立研究员)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 CircuitProbe通过激活统计预测Transformer中的推理电路,利用稳定性区和幅度区检测,提升推理效率,验证了其在多种模型上的有效性。

Comments 11 pages, 1 figure, 3 tables. Code available at https://github.com/agenticclass/circuitprobe

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28610 2026-04-01 cs.CV cs.AI cs.CL 81%

ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning

ResAdapt:面向高效多模态推理的自适应分辨率

Huanxuan Liao, Zhongtao Jiang, Yupu Hao, Yuqiao Tan, Shizhu He, Ben Wang, Jun Zhao, Kun Xu, Kang Liu

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 ResAdapt通过自适应输入分辨率框架,在保持高空间分辨率的同时提升多模态推理效率,尤其在压缩条件下显著提升性能。

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22715 2026-04-01 cs.CV cs.AI cs.CL cs.MM 81%

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

ReAG:基于推理的生成用于基于知识的视觉问答

Alberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi, Davide Caffagni, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳大学和雷焦艾米利亚大学) University of Pisa(比萨大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 ReAG通过结合粗粒度和细粒度检索及批评模型过滤无关信息,提升知识密集型视觉问答的准确性和可解释性。

Comments CVPR 2026 - Project page: https://aimagelab.github.io/ReAG/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12476 2026-03-27 cs.CL cs.LG 81%

Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation

基于检索-推理的大型语言模型合成临床试验生成

Zerui Xu, Fang Wu, Yingzhou Lu, Yuanyuan Zhang, Yue Zhao

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 国家信息与自动化研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒实验室) University of Chicago(芝加哥大学) Stanford University(斯坦福大学) Purdue University(普渡大学) University of Southern California(南加州大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出基于检索-推理框架的合成临床试验生成方法,利用LLM生成标注二元结果的合成试验报告,通过检索模块和推理模块提升生成质量,实验证明合成数据可有效增强真实数据集并提升临床试验预测性能。

Comments Published in ACM BCB 2025. 9 pages, 4 figures, 5 tables (Main paper + Supplementary Materials)

Journal ref Proceedings of the 16th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM BCB 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22352 2026-03-25 cs.LG cs.AI 81%

WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement

WIST:基于网页的迭代自我对战树用于领域针对性推理改进

Fangyuan Li, Pengfei Li, Shijie Wang, Junqi Gao, Jianxing Liu, Biqing Qi, Yuqiang Li

机构 * Harbin Institute of Technology(哈尔滨工业大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 WIST通过基于网页的迭代自我对战树框架,无需预设领域语料,直接从开放网络学习,提升领域推理能力,优于纯内生自我进化和语料基础自我对战基线。

Comments 23 pages, 4 figures. Submitted to ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22117 2026-03-24 cs.LG cs.AI 81%

On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation

关于RLVR更新方向的LLM推理:识别与利用

Kexin Huang, Haoming Meng, Junkang Wu, Jinda Lu, Chiyu Ma, Ziqian Chen, Xue Wang, Bolin Ding, Jiancan Wu, Xiang Wang, Xiangnan He, Guoyin Wang, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文探讨RLVR更新方向对LLM推理的影响,提出通过Δlogp识别关键更新,并应用于测试时插值和训练时重加权以提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10679 2026-03-24 cs.AI cs.LG 81%

Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models

您的推理模型是推理还是猜测?对分层推理模型的机理分析

Zirui Ren, Ziming Liu

机构 * Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究院) Department of Physics, Tsinghua University, Beijing, China(清华大学物理系) College of AI, Tsinghua University, Beijing, China(清华大学人工智能学院)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 研究揭示分层推理模型在简单谜题中易失效,存在猜测动态和多固定点问题,提出数据增强、输入扰动和模型自举策略提升Sudoku-Extreme准确率至96.9%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16553 2026-03-18 cs.CL cs.AI 81%

EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models

EmoLLM:基于评估的认知-情感共推理在大语言模型中

Yifei Zhang, Mingyang Li, Henry Gao, Liang Zhao

机构 * Department of Computer Science, Emory University(埃默里大学计算机科学系)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 EmoLLM通过引入评估推理图结构,实现认知与情感的协同推理,提升对话中情感状态和响应质量,同时保持事实可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15547 2026-03-17 cs.CL cs.AI cs.HC 81%

Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation

大语言模型能否建模错误学生推理?一种关于干扰项生成的案例研究

Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan

机构 * ETH Zürich, Switzerland(苏黎世联邦理工学院,瑞士) Bocconi University, Italy(博科尼大学,意大利) University of Central Florida, USA(中央佛罗里达大学,美国)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文研究了大语言模型在生成多项选择题干扰项时对错误推理的建模能力,发现模型通常先正确解决问题,再模拟可能的误解,最后选择干扰项,且正确解的提示能提升干扰项质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11248 2026-03-17 cs.LG cs.CL 81%

Reasoning-Grounded Natural Language Explanations for Language Models

基于推理的自然语言解释方法用于语言模型

Vojtech Cahlik, Rodrigo Alves, Pavel Kordik

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出一种通过推理过程生成可信自然语言解释的方法,通过将推理过程转化为token序列并解码为自然语言,提升解释的准确性与答案质量。

Comments (v2) Added acknowledgements section

Journal ref In: Guidotti, R., Schmid, U., Longo, L. (eds) Explainable Artificial Intelligence. xAI 2025. Communications in Computer and Information Science, vol 2578. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02025 2026-03-03 cs.LG cs.AI 81%

Revealing Combinatorial Reasoning of GNNs via Graph Concept Bottleneck Layer

通过图概念瓶颈层揭示图神经网络的组合推理

Yue Niu, Zhaokai Sun, Jiayi Yang, Xiaofeng Cao, Rui Fan, Xin Sun, Hanli Wang, Wei Ye

机构 * Tongji University, Shanghai, China(同济大学) City University of Macau, Macau, China(澳门城市大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出图概念瓶颈层,通过将概念视为图词并利用语言模型学习嵌入,提升GNNs的组合推理可解释性与性能。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01326 2026-03-03 cs.CL cs.LG 81%

Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning

真相作为轨迹:内部表示揭示大语言模型推理的本质

Hamed Damirchi, Ignacio Meza De la Jara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang, Javen Shi

机构 * Australian Institute for Machine Learning, Adelaide University(澳大利亚机器学习研究所,阿德莱德大学) Monash University(莫纳什大学) Concordia University(康科迪亚大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 TaT通过分析大语言模型推理过程中的层间几何位移,揭示有效推理与虚假行为的区别,提升模型可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19473 2026-03-03 cs.LG cs.AI 81%

WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning

WavefrontDiffusion: 动态解码调度以提升推理性能

Haojin Yang, Rui Hu, Zequn Sun, Rui Zhou, Yujun Cai, Yiwei Wang

机构 * School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室) The University of Queensland(昆士兰大学) University of California, Merced(加州大学梅尔德分校)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 WavefrontDiffusion通过动态解码调度提升推理和代码生成的性能与语义连贯性。

Comments 19 pages. 3 figures

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22611 2026-03-03 cs.LG cs.AI 81%

Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning

分位数优势估计:稳定LLM推理的RLVR

Junkang Wu, Kexin Huang, Jiancan Wu, An Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 通过分位数优势估计方法,稳定RLVR训练过程,提升LLM推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21413 2026-03-03 cs.CL cs.AI 81%

RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning

RefTool: 基于参考的工具创建用于知识密集型推理

Xiao Liu, Da Yin, Zirui Wu, Yansong Feng

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王萱计算机技术研究所) University of Chicago(芝加哥大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他推理 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 RefTool通过基于参考的工具创建框架,提升LLMs在知识密集型任务中的推理能力,实现更准确和高效的工具生成与应用。

Comments Accepted by ICLR 2026. Code is available at https://github.com/xxxiaol/RefTool

详情

展开后加载摘要…

URL PDF HTML 收藏