arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10364 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10364 篇

2510.06243 2026-01-21 cs.CL cs.AI 90%

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

CoT Referring: 通过 grounded 推理改进指称表达任务

Qihua Dong, Luis Figueroa, Handong Zhao, Kushal Kafle, Jason Kuen, Zhihong Ding, Scott Cohen, Yun Fu

机构 * Adobe Research(Adobe研究院) Northeastern University(东北大学)

专题命中 推理评测 :reasoning(title,abstract);CoT(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 通过 grounded 推理改进指称表达任务,提出CoT Referring方法,提升多模态大语言模型在复杂指称场景中的性能。

Comments MLLM, Referring Expression Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06289 2026-01-13 cs.CL cs.LG 90%

How well can off-the-shelf LLMs elucidate molecular structures from mass spectra using chain-of-thought reasoning?

离线大语言模型如何通过链式推理解析质谱中的分子结构?

Yufeng Wang, Lu Wei, Lin Liu, Hao Xu, Haibin Ling

机构 * Stony Brook University(石溪大学) Stanford University(斯坦福大学) Harvard Medical School(哈佛医学院)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.LG

AI总结 本研究评估了离线大语言模型通过链式推理解析质谱数据的能力,发现其在化学准确性上存在局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05871 2025-10-08 cs.AI cs.LG 90%

Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering

Josefa Lia Stoisser, Lawrence Phillips, Aditya Misra, Tom A. Lamb, Philip Torr, Marc Boubnovski Martell, Julien Fauqueur, Kaspar Märtens

机构 * Novo Nordisk(诺华制药) University of Oxford(牛津大学)

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title);chain-of-thought(abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00785 2025-09-10 cs.AI cs.CV cs.LG 90%

GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning

Sahiti Yerramilli, Nilay Pande, Rynaa Grover, Jayant Sravan Tamarapalli

机构 * Google(谷歌) Waymo

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05518 2025-06-30 cs.CL cs.AI 90%

Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

James Chua, Edward Rees, Hunar Batra, Samuel R. Bowman, Julian Michael, Ethan Perez, Miles Turpin

机构 * Speechmatics, Apollo Research(Speechmatics与Apollo研究机构) University of Oxford(牛津大学) NYU, Anthropic(纽约大学与Anthropic) NYU(纽约大学) Anthropic, NYU(Anthropic与纽约大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10937 2025-05-19 cs.CL cs.AI 90%

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations

Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang, Xiangzhong Fang

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Cloud Computing(阿里云计算)

专题命中 推理评测 :reasoning(title,abstract);CoT(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07956 2025-04-11 cs.CV cs.AI cs.CL 90%

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Yukun Qi, Yiming Zhao, Yu Zeng, Xikun Bao, Wenxuan Huang, Lin Chen, Zehui Chen, Jie Zhao, Zhongang Qi, Feng Zhao

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11154 2025-03-17 cs.CL cs.AI 90%

Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models

Shaotian Yan, Chen Shen, Wenxiao Wang, Liang Xie, Junjie Liu, Jieping Ye

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted by ICLR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12025 2025-02-18 cs.AI cs.CL 90%

SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, Radha Poovendran

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14794 2024-11-25 cs.CV cs.AI cs.CL 90%

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Songhao Han, Wei Huang, Hairong Shi, Le Zhuo, Xiu Su, Shifeng Zhang, Xu Zhou, Xiaojuan Qi, Yue Liao, Si Liu

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09220 2024-10-15 cs.CL cs.CY cs.LG 90%

M3Hop-CoT: Misogynous Meme Identification with Multimodal Multi-hop Chain-of-Thought

Gitanjali Kumari, Kirtan Jain, Asif Ekbal

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(title,abstract);reasoning(abstract);分类 cs.CL、cs.LG

Comments 34 Pages. Accepted in The 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024). Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18344 2024-10-15 cs.CL cs.AI 90%

Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning

Jiachun Li, Pengfei Cao, Chenhao Wang, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, Jun Zhao

专题命中 推理评测 :reasoning(title,abstract);CoT(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments Accepted as a long paper to ACL 2024 Main, 25 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16473 2024-05-28 cs.CV cs.AI cs.CL 90%

M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, Wanxiang Che

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted at ACL2024 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18564 2024-04-30 cs.CL cs.AI 90%

Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning

Wen-Yu Chang, Yun-Nung Chen

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2308.14266

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04461 2024-03-21 cs.CL cs.CV cs.LG 90%

Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models

Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, Ajay Divakaran

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.LG

Comments NAACL 2024 Main Conference. The data is released at https://github.com/Yangyi-Chen/CoTConsistency

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13552 2023-10-24 cs.CL cs.AI 90%

Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning

Jinyuan Wang, Junlong Li, Hai Zhao

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments Accepted by Findings of EMNLP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10740 2026-06-16 cs.AI cs.CL cs.LG 新提交 89%

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

当思维链更清楚时:多轮推理模型的失败模式

Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi

专题命中 推理评测 :CoT(summary_cn,abstract);reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出CoT-Output 2x2安全矩阵诊断多轮推理模型隐藏的时间动态失败,发现监督悖论和上下文注入失败两种可复现漏洞。

Comments Accepted at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07355 2026-05-29 cs.MM cs.SD 89%

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

AV-EMO-Reasoning: 在具有视听线索的全模态大语言模型中基准测试情感推理能力

Dingkun Zhou, Krish Patel, Ajay Kankipati, Akshaj Gupta, Zeyi Austin Li, Mohul Shukla, Vibhor Narang, Sara Kofman, Zongli Ye, Grace Wang, Xiaoyu Shi, Tingle Li, Guan-Ting Lin, Kan Jen Cheng, Huang-Cheng Chou, Jiachen Lian, Gopala Anumanchipalli

机构 * UC Berkeley(加州大学伯克利分校) South China University of Technology(华南理工大学) Zhejiang University(浙江大学) National Taiwan University(台湾大学) University of Southern California(美国南加州大学)

专题命中 推理评测 :reasoning(title,title_cn)

AI总结 提出AV-EMO-Reasoning基准,通过合成和真实世界的视听对话数据集及情感感知与交互推理指标,系统评估全模态大语言模型的情感推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06206 2026-05-21 cs.RO cs.CV 89%

Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

Affordance-R1: 为多模态大语言模型中的通用化 affordance 推理设计的强化学习

Hanqing Wang, Shaoyang Wang, Yiming Zhong, Zemin Yang, Jiamin Wang, Zhiqing Cui, Jiahao Yuan, Yifan Han, Mingyu Liu, Yuexin Ma

机构 * The Hong Kong University of Science and Technology (GZ)(香港科技大学(广州)) National University of Singapore(新加坡国立大学) ShanghaiTech University(上海科技大学) East China Normal University(华东师范大学) Nanjing University of Information Science & Technology(南京信息工程大学) Zhejiang University(浙江大学) Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 推理评测 :CoT(summary_cn,abstract);reasoning(title,abstract);chain-of-thought(abstract)

AI总结 本文提出 Affordance-R1,一种结合认知 CoT 引导的 Group Relative Policy Optimization (GRPO) 的统一 affordance 地标框架,通过强化学习实现零样本泛化和测试时推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18956 2026-05-20 cs.CV 89%

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

MotionMERGE: 一种用于人体动作编辑、推理、生成和解释的多粒度框架

Bizhu Wu, Jinheng Xie, Wenting Chen, Zhe Kong, Jianfeng Ren, Linlin Shen, Ruibin Bai, Rong Qu

机构 * Computer Vision Institute, School of Computer Science and Software Engineering, Shenzhen University(计算机视觉研究院,计算机科学与软件工程学院,深圳大学) Guangdong Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室,深圳大学) School of Computer Science, University of Nottingham Ningbo China(Nottingham Ningbo 中国计算机科学学院) Department of Electrical and Computer Engineering, National University of Singapore(电子与计算机工程系,新加坡国立大学) Department of Radiation Oncology, Stanford University(放射肿瘤科,斯坦福大学) Sun Yat-sen University(中山大学) School of Computer Science, University of Nottingham(计算机科学学院,Nottingham大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract,abstract_cn);CoT(abstract,abstract_cn)

AI总结 本文提出MotionMERGE框架,通过细粒度语言引导的动作控制、跨粒度协同预训练和细粒度动作-语言对齐,实现了更精确的动作生成、理解和编辑,并建立了新的细粒度文本驱动动作编辑和动作引导推理基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04973 2024-07-09 cs.AI cs.CL cs.CV cs.LG 89%

LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Yijia Xiao, Edward Sun, Tianyu Liu, Wei Wang

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments LogicVista benchmarks the logical reasoning of multimodal large language models in visual tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11281 2026-04-21 cs.CL cs.AI cs.CY 89%

ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection

ToxiFrench:通过CoT微调提升法语毒性检测的语言模型基准测试

Axel Delaval, Shujian Yang, Haicheng Wang, Han Qiu, Jialiang Lu

机构 * École Polytechnique(巴黎高等理工学院) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 推理评测 :CoT(title,title_cn);chain-of-thought(abstract,comments);分类 cs.CL、cs.AI

AI总结 本文提出ToxiFrench数据集,通过半自动化标注流程构建,发现小语言模型在毒性检测任务中表现更优,并提出动态加权损失策略提升模型忠实度,Qwen3-4B模型在基准测试中取得最佳性能。

Comments 22 pages, 5 figures, 11 tables. This paper introduces TOXIFRENCH, a benchmark of 53,622 comments for French toxicity detection. It proposes a Chain-of-Thought fine-tuning method with a dynamic weighted loss. The fine-tuned 4B model (Qwen3-4B) achieves state-of-the-art performance, outperforming larger models like GPT-4o and DeepSeek-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04852 2026-04-07 cs.CR cs.AI 89%

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework

通过结构化提示框架强化LLMs中以人为中心的链式推理完整性

Jiling Zhou, Aisvarya Adeseye, Seppo Virtanen, Antti Hakkala, Jouni Isoaho

机构 * Department of Computing, University of Turku(图尔库大学计算机系)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI

AI总结 本文提出结构化提示框架以提升LLMs链式推理的可靠性及安全威胁检测能力,通过16个核心维度因素优化推理过程,实验证明在DDoS攻击检测中推理能力提升达40%,并获得高人评一致性。

Comments This paper has been accepted at the 12th Intelligent Systems Conference (IntelliSys 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01702 2026-04-07 cs.CL 89%

On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning

在长链式思维监督微调中推理模式在泛化差异中的作用

Zhaoyi Li, Xiangyu Xi, Zhengyu Chen, Wei Wang, Gangwei Jiang, Ranran Shen, Linqi Song, Ying Wei, Defu Lian

机构 * University of Science and Technology of China(中国科学技术大学) Meituan LongCat Team(美团龙猫团队) City University of Hong Kong(香港城市大学) Zhejiang University(浙江大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 本文研究不同来源的长链式思维轨迹对模型泛化性能的影响,发现训练损失低并不一定带来更好的泛化能力,提出过滤频繁分支轨迹以提升SFT泛化性能。

Comments Under Review. version2: correct typos in Table 4 and add an ablation study (Table 5)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05618 2026-03-09 cs.CL 89%

Safer Reasoning Traces: Measuring and Mitigating Chain-of-Thought Leakage in LLMs

更安全的推理轨迹:衡量和减轻大语言模型中的推理链泄露

Patrick Ahrend, Tobias Eder, Xiyang Yang, Zhiyi Pan, Georg Groh

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 本文研究了大语言模型中推理链泄露问题,提出了一种模型无关的框架来衡量和减轻隐私风险,通过比较不同模型和预算下的泄露情况,提出了混合门卫策略以平衡效用与风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04040 2026-03-03 cs.AI 89%

FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning

FaithCoT-Bench: 对链式推理实例级忠实度的基准测试

Xu Shen, Song Wang, Zhen Tan, Laura Yao, Xinyu Zhao, Kaidi Xu, Xin Wang, Tianlong Chen

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI

AI总结 FaithCoT-Bench提出了一种统一的基准,用于评估链式推理在实例级的忠实度,通过专家标注的轨迹和系统评估揭示了现有方法的优缺点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04230 2026-01-14 cs.CL 89%

Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

通过语言混合的推理模型推动多语言推理

Guijin Son, Donghun Yang, Hitesh Laxmichand Patel, Amit Agarwal, Hyunwoo Ko, Chanuk Lim, Srikant Panda, Minhyuk Kim, Nikunj Drolia, Dasol Choi, Kyong-Ha Lee, Youngjae Yu

机构 * OneLineAI KISTI Oracle AI Korea University(韩国大学) Modulabs Seoul National University(首尔国立大学) University College Dublin(都柏林大学学院) Yonsei University(延世大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 本文提出语言混合的推理方法,通过多语言推理提升模型性能,展示在韩国案例研究中达到最先进的表现。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05076 2026-01-09 cs.AI 89%

Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models

链式净化思维:在大推理模型的CoT中消除PII泄露

Arghyadeep Das, Sai Sreenivas Chintha, Rishiraj Girmal, Kinjal Pandey, Sharvi Endait

专题命中 推理评测 :reasoning(title,abstract);CoT(title,abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出PII-CoT-Bench,通过部署干预措施减少大推理模型中链式思维推理的PII泄露,证明隐私优先推理可在不牺牲性能的情况下实现。

Comments 12 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13127 2025-12-03 cs.CL 89%

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

通过监督的链式思维推理促进长上下文理解

Jingyang Lin, Andy Wong, Tian Xia, Shenghua He, Hui Wei, Mei Han, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) PAII Inc.(PAII公司) University of California, Merced(加州梅尔德大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 本文提出基于属性的代理推理框架PAI,通过生成包含中间推理步骤的合成数据集LongFinanceQA,提升LLMs在金融领域的长上下文理解能力。

Comments Main Conference of EMNLP 2025, Project Page: https://long-pai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13771 2025-11-19 cs.CR cs.AI 89%

ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning

Shaowei Guan, Yu Zhai, Zhengyu Zhang, Yanze Wang, Hin Chi Kwok

机构 * Centre for Smart Health, School of Nursing, The Hong Kong Polytechnic University(智能健康研究中心、护理学院、香港理工大学) Department of Language Science and Technology, The Hong Kong Polytechnic University(语言科学与技术系、香港理工大学) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系、香港理工大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏