arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-04-08 至 2026-04-08 共收录 7 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 7 篇

2604.05868 2026-04-08 cs.CL 79%

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models

理解并行采样与顺序采样在大推理模型中的性能差距

Xiangming Gu, Soham De, Larisa Markeeva, Petar Veličković, Razvan Pascanu

机构 * Google DeepMind(谷歌DeepMind) National University of Singapore(新加坡国立大学)

专题命中 测试时计算 :reasoning(title,abstract);分类 cs.CL

AI总结 本文研究大推理模型中并行与顺序采样性能差异,发现并行采样表现更优,提出探索不足是主要原因。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05939 2026-04-08 cs.AI cs.HC 70%

Context-Value-Action Architecture for Value-Driven Large Language Model Agents

基于价值的大型语言模型代理的上下文-价值-行动架构

TianZe Zhang, Sirui Sun, Yuhang Xie, Xin Zhang, Zhiqiang Wu, Guojie Song

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) Yuanpei College, Peking University(北京大学元培学院) School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院) Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学机器感知重点实验室(教育部)) PKU-Wuhan Institute for Artificial Intelligence(北大武汉人工智能研究院)

专题命中 测试时计算 :reasoning(abstract);verifier(abstract);分类 cs.AI

AI总结 本文提出CVA架构,通过价值验证器缓解价值极化,提升代理行为的准确性和可解释性。

Comments Accepted to Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05716 2026-04-08 cs.AI 70%

Can Large Language Models Reinvent Foundational Algorithms?

大语言模型能否重新发明基础算法?

Jian Zhao, Haoren Luo, Yu Wang, Yuhan Cao, Pingyue Sheng, Tianxing He

机构 * Xiongan AI Institute(雄安人工智能研究院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Shanghai Qi Zhi Institute(上海期智研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 测试时计算 :reasoning(abstract);verifier(abstract);分类 cs.AI

AI总结 本文研究大语言模型能否在受控环境中重新发明基础算法,通过Unlearn-and-Reinvent流程验证模型在无提示、低提示和高提示下的表现,发现生成验证器在推理过程中起到关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06013 2026-04-08 cs.AI cs.CL 62%

Epistemic Blinding: An Inference-Time Protocol for Auditing Prior Contamination in LLM-Assisted Analysis

知识盲区:一种用于审计LLM辅助分析中先验污染的推理时协议

Michael Cuccarese

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种在代理系统中用于审计LLM辅助分析中先验污染的推理时协议,通过替换实体标识符以匿名代码并比较输出与未盲对照组,以恢复审计性。

Comments code and LLM skill at: https://github.com/mcuccarese/epistemic-blinding 7 pages 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05250 2026-04-08 cs.LG cs.CL 62%

DualDiffusion: A Speculative Decoding Strategy for Masked Diffusion Models

DualDiffusion: 一种用于掩码扩散模型的推测解码策略

Satyam Goyal, Kushal Patel, Tanush Mittal, Arjun Laxman

机构 * University of Michigan(密歇根大学)

专题命中 测试时计算 :verifier(abstract);分类 cs.CL、cs.LG

AI总结 DualDiffusion结合高效近似模型与准确验证模型,通过多步轻量级生成后单步验证,实现生成步骤与准确性之间的更优帕累托前沿。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04987 2026-04-08 cs.LG cs.AI math.OC stat.ML 62%

Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling

Cactus:通过受约束的接受推测采样加速自动回归解码

Yongchang Hao, Lili Mou

机构 * Dept. Computing Science & Alberta Machine Intelligence Institute (Amii), University of Alberta(阿尔伯塔大学计算机科学系与阿尔伯塔机器智能研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能讲席)

专题命中 测试时计算 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 本文提出Cactus方法,通过受约束优化框架,在保证与验证器分布可控差异的同时提升接受率,有效解决传统推测采样在分布匹配上的限制问题。

Comments Camera-ready version. Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18740 2026-04-08 cs.IR 50%

Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation

具有自适应偏好优化的多模态大语言模型用于序列推荐

Yu Wang, Yonghui Yang, Le Wu, Yi Zhang, Fei Liu, Richang Hong

专题命中 测试时计算 :reasoning(abstract)

AI总结 本文提出HaNoRec框架,通过自适应偏好优化解决多模态序列推荐中的样本不平衡和跨模态语义偏差问题,提升推荐性能。

Comments Accepted by SIGIR 2026 (Full Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏