arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1990 信号源:cs.CL, cs.AI, cs.LG

1. 测试时计算 1990 篇

2412.12564 2025-06-25 cs.CL 70%

Evaluating Zero-Shot Multilingual Aspect-Based Sentiment Analysis with Large Language Models

Chengyan Wu, Bolei Ma, Zheyu Zhang, Ningyuan Deng, Yanqing He, Yun Xue

机构 * Guangdong Provincial Key Laboratory of Quantum Engineering and Quantum Materials, School of Electronic Science and Engineering (School of Microelectronics)(广东量子工程与材料重点实验室,电子科学与工程学院(微电子学院)) South China Normal University(华南师范大学) Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东智能信息处理重点实验室) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Technical University of Munich(慕尼黑技术大学)

专题命中 测试时计算 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Preprint; Paper accepted at International Journal of Machine Learning and Cybernetics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24857 2025-06-02 cs.LG 70%

Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking

Heli Ben-Hamu, Itai Gat, Daniel Severo, Niklas Nolte, Brian Karrer

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 测试时计算 :reasoning(abstract);math reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03261 2025-05-20 cs.CL 70%

Can Frontier LLMs Replace Annotators in Biomedical Text Mining? Analyzing Challenges and Exploring Solutions

Yichong Zhao, Susumu Goto

机构 * The University of Tokyo(东京大学) Graduate School of Frontier Sciences(前沿科学研究生院) Department of Computational Biology and Medical Sciences(计算生物学与医学科学系) Database Center for Life Science(生命科学数据库中心) Joint Support-Center for Data Science Research(数据科学研究联合支持中心)

专题命中 测试时计算 :reasoning(abstract);test-time compute(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10772 2025-05-19 cs.CL 70%

Ranked Voting based Self-Consistency of Large Language Models

Weiqin Wang, Yile Wang, Hui Huang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01828 2025-05-05 cs.RO cs.LG 70%

From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment

Yilin Wu, Ran Tian, Gokul Swamy, Andrea Bajcsy

机构 * Carnegie Mellon University(卡内基梅隆大学) UC Berkeley(伯克利大学)

专题命中 测试时计算 :reasoning(abstract);verifier(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06251 2025-04-29 cs.AI 70%

Quasi-random Multi-Sample Inference for Large Language Models

Aditya Parashar, Aditya Vikram Singh, Avinash Amballa, Jinlin Lai, Benjamin Rozonoyer

机构 * College of Information & Computer Sciences, University of Massachusetts Amherst(信息与计算机科学学院,马萨诸塞大学阿默斯特分校)

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16760 2025-04-24 cs.AI 70%

Lightweight Latent Verifiers for Efficient Meta-Generation Strategies

Bartosz Piotrowski, Witold Drzewakowski, Konrad Staniszewski, Piotr Miłoś

机构 * IDEAS NCBR &Witold Drzewakowski University of Warsaw(IDEAS NCBR与华沙大学) IDEAS NCBR &Konrad Staniszewski University of Warsaw(IDEAS NCBR与华沙大学) IDEAS NCBR NVIDIA &Piotr Miłoś IMPAN(IDEAS NCBR NVIDIA与波兰科学院)

专题命中 测试时计算 :reasoning(abstract);self-correction(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13592 2025-04-22 cs.CL 70%

Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling

Zihao Feng, Xiaoxue Wang, Ziwei Bai, Donghang Su, Bowen Wu, Qun Yu, Baoxun Wang

专题命中 测试时计算 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13007 2025-02-20 cs.CL 70%

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament

Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, Juanzi Li

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments in progress work

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18585 2025-02-19 cs.CL 70%

Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, Dong Yu

专题命中 测试时计算 :reasoning(abstract);test-time compute(abstract);分类 cs.CL

Comments 1. We have updated the results for DeepSeek-R1, and all of our original conclusions remain valid. 2. Our proposed Tip approach remains effective in Best-of-N scenarios (e.g., self-consistency and Laconic Decoding) when built on DeepSeek-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11157 2025-02-18 cs.AI 70%

Dyve: Thinking Fast and Slow for Dynamic Process Verification

Jianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen, Qiang Xu

专题命中 测试时计算 :reasoning(abstract);verifier(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14427 2025-02-13 cs.LG 70%

GraphSOS: Graph Sampling and Order Selection to Help LLMs Understand Graphs Better

Xu Chu, Hanlin Xue, Zhijie Tan, Bingce Wang, Tong Mo, Weiping Li

专题命中 测试时计算 :reasoning(abstract);CoT(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18841 2025-02-03 cs.LG cs.CR 70%

Trading Inference-Time Compute for Adversarial Robustness

Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak, Stephanie Lin, Sam Toyer, Yaodong Yu, Rachel Dias, Eric Wallace, Kai Xiao, Johannes Heidecke, Amelia Glaese

专题命中 测试时计算 :reasoning(abstract);test-time compute(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11514 2025-01-16 cs.CL 70%

Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Yi Fang, Moxin Li, Wenjie Wang, Hui Lin, Fuli Feng

专题命中 测试时计算 :reasoning(abstract);self-correction(abstract);分类 cs.CL

Comments accepted by COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08523 2024-09-17 cs.CL 70%

Eir: Thai Medical Large Language Models

Yutthakorn Thiprak, Rungtam Ngodngamthaweesuk, Songtam Ngodngamtaweesuk

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

Comments typos corrected, and references added

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04361 2024-04-09 cs.CL 70%

Deciphering Political Entity Sentiment in News with Large Language Models: Zero-Shot and Few-Shot Strategies

Alapan Kuila, Sudeshna Sarkar

专题命中 测试时计算 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Accepted in PoliticalNLP workshop co-located with LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14479 2023-10-24 cs.CL 70%

DetectGPT-SC: Improving Detection of Text Generated by Large Language Models through Self-Consistency with Masked Predictions

Rongsheng Wang, Qi Li, Sihong Xie

专题命中 测试时计算 :reasoning(abstract);logical reasoning(abstract);分类 cs.CL

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11610 2022-10-26 cs.CL 70%

Large Language Models Can Self-Improve

Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, Jiawei Han

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00747 2022-07-05 cs.CL 70%

Rationale-Augmented Ensembles in Language Models

Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Denny Zhou

专题命中 测试时计算 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11731 2026-02-10 cs.LG cs.AI cs.CL 69%

Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models

Dist2ill: 基于分布的蒸馏用于大语言模型中单次传递的不确定性估计

Yicong Zhao, King Yeung Tsang, Harshil Vejendla, Haizhou Shi, Zhuohang Li, Zhigang Hua, Qi Xu, Tunyu Zhang, Yi Wang, Ligong Han, Bradley A. Malin, Hao Wang

机构 * Rutgers University(罗格斯大学) Vanderbilt University(范德比尔特大学) Meta

专题命中 测试时计算 :reasoning(abstract,comments);分类 cs.CL、cs.AI、cs.LG

AI总结 Dist2ill通过在单次推断中生成多个多样化的推理路径,利用轻量级模块近似置信度分数,实现大语言模型中更准确的不确定性估计。

Comments Preprint; work in progress. Update Log: 05/2025 (v1&v2): Introduced Dist2ill (previously named EUD) for efficient uncertainty estimation, focusing on discriminative reasoning tasks. 02/2026 (v3): Extended Dist2ill to a unified framework supporting both discriminative and generative reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15709 2026-08-18 cs.DB 新提交 67%

Logos: Certified Order-Sensitive SQL Rewrites with Mechanized Semantics and LLM Guidance

Logos:结合机械化语义与LLM指导的可验证顺序敏感型SQL重写

Jingyu Ke, Jingyang Li, Guoqiang Li

专题命中 测试时计算 :reasoning(abstract);verifier(abstract)

AI总结 Logos是一款结合Rocq机械化语义与LLM指导的SQL重写验证工具,解决了现有工具的不足,在389个查询对的评估中解决率达86.9%,优于基线工具SQLSolver。

Comments 13 pages, 4 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12307 2026-08-13 cs.LG cs.AI cs.CL 新提交 67%

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

测试时的AI辅助AI:通过 harness 实现强模型到弱模型的能力迁移

Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

机构 * Salesforce AI Research(Salesforce人工智能研究院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出在测试时通过构建 harness,无需重新训练即可将强模型的认知结构迁移给弱模型,使目标模型在四个心理理论基准上的平均性能从0.49提升至0.91,为模型能力迁移提供了新的思路。

Comments 23 Pages, 12 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00715 2026-08-11 cs.CL cs.AI cs.LG 版本更新 67%

To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval

记忆还是检索:考虑RAG的缩放规律

Karan Singh, Michael Yu, Varun Gangal, Zhuofu Tao, Sachin Kumar, Emmy Liu, Steven Y. Feng

机构 * Stanford University(斯坦福大学) Independent Researcher(独立研究员) Patronus AI The Ohio State University(俄亥俄州立大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨了预训练知识与检索知识的平衡,提出三维缩放框架,揭示在不同模型规模和任务类型下检索的边际效用,为语言模型设计提供数据资源分配指导。

Comments Code available at https://github.com/DegenAI-Labs/RAG-Scaling-Laws

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00066 2026-08-04 cs.CV 新提交 67%

PhysAgent: A Multi-Agent Framework for Reliable Remote Heart Rate Estimation

PhysAgent:用于可靠远程心率估计的多智能体框架

Yehui Yang, Bo Zhao, Junzhe Cao, Hui Ma, Yue Sun, Wenjin Wang, Zitong Yu

专题命中 测试时计算 :reasoning(abstract);verifier(abstract)

AI总结 PhysAgent是一种推理时多智能体候选验证框架,以多个基础rPPG估计器输出为待验证生理假设,结合Qwen3-VL-4B多模态大语言模型推理与确定性融合,提升远程心率估计的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23993 2026-08-04 cs.CL cs.AI cs.DB cs.LG cs.MA 67%

EPM-RL: Reinforcement Learning for On-Premise Product Mapping in E-Commerce

EPM-RL:电商产品映射的强化学习方法

Minhyeong Yu, Wonduk Seo

机构 * AI Research, Enhans(AI研究,Enhans)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出EPM-RL框架,通过强化学习提升电商产品映射的准确性和效率,解决传统方法依赖外部API和高成本的问题,实现私有部署和低成本运营。

Comments preprint

Journal ref Proceedings of the AAAI Symposium Series, vol. 9, no. 1, p. 227, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29596 2026-08-03 cs.RO cs.CV 新提交 67%

FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

FibVLA:一种采用斐波那契采样的高效时序视觉-语言-动作模型

Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang

专题命中 测试时计算 :reasoning(abstract);planning(abstract)

AI总结 本文提出FibVLA框架,通过对数事后采样和流匹配等技术,解决了视觉-语言-动作模型时序信息捕捉与推理效率的矛盾,提升了动作性能与实时响应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01551 2026-07-27 cs.LG cs.AI cs.CL 版本更新 67%

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

基于重新定义的逐步优势的自引导过程奖励优化用于过程强化学习

Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong

机构 * Terminus Group(Terminus集团) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对过程强化学习中引入奖励模型计算开销大且缺乏统一优势估计框架的问题,提出SPRO框架,通过从策略模型导出奖励及重新定义优势进行过程感知强化学习,实验显示其训练效率高、准确率提升且无额外开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24927 2026-07-22 cs.CL cs.AI cs.LG 版本更新 67%

Large Language Models Explore by Latent Distilling

通过潜在蒸馏探索大语言模型

Yuanhao Zeng, Ao Lu, Lufei Li, Zheng Zhang, Yexin Li, Kan Ren

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China(人工智能通用基础理论国家重点实验室,BIGAI,北京,中国) School of Information Science and Technology, ShanghaiTech University, Shanghai, China(信息科学与技术学院,上海交通大学,上海,中国)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出Exploratory Sampling方法,通过蒸馏模型引导生成过程,提升语义多样性与推理效率,在数学、科学和代码生成中表现优异。

Comments 25 pages, 5 figures. Accepted in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11388 2026-07-14 cs.AI cs.CL cs.LG cs.MA 新提交 67%

StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure

StructAgent:利用统一因果结构驾驭长时程数字智能体

Wenyi Wu, Sibo Zhu, Kun Zhou, Aayush Salvi, Zixuan Song, Biwei Huang

机构 * University of California, San Diego(加利福尼亚大学圣地亚哥分校) Aether AI Lab(以太人工智能实验室)

专题命中 测试时计算 :verifier(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究针对长时程任务中智能体任务进展难解释、验证和恢复的问题,提出StructAgent框架,通过统一因果结构维护任务进展并规范工作流程,实验证明其能提升多种模型性能且具有通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09438 2026-07-13 cs.CL cs.AI cs.LG 新提交 67%

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

多语言视觉多项选择题中小视觉语言模型的测试时缩放

Spiros Baxevanakis, Peng-Jian Yang

机构 * ImageCLEF Lab(图像CLEF实验室) University of Amsterdam(阿姆斯特丹大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究小型视觉语言模型在多语言视觉多项选择题中的测试时缩放,比较多种方法,发现运行条件重要,可解析性是关键,增加解码预算有帮助,复杂方法贡献小,最佳配置在测试集上表现出色,位居排行榜首位。

Comments 14 pages, 2 figures, accepted at ImageCLEF 2026

详情

展开后加载摘要…

URL PDF HTML 收藏