arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10379 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10379 篇

2509.06409 2025-09-09 cs.AI 88%

Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning

Yihong Luo, Wenwu He, Zhuo-Xu Cui, Dong Liang

机构 * Fujian University of Technology(福建工程学院) Fujian Provincial Key Laboratory of Big Data Mining and Applications(福建省大数据挖掘与应用重点实验室) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(生物医学成像科学与系统重点实验室,中国科学院)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15727 2025-07-01 cs.NI cs.AI cs.CR cs.IR 88%

Retrieval Augmented Generation Based LLM Evaluation For Protocol State Machine Inference With Chain-of-Thought Reasoning

Youssef Maklad, Fares Wael, Wael Elsersy, Ali Hamdi

机构 * Faculty of Computer Science, MSA University(计算机科学学院,MSA大学)

专题命中 推理评测 :chain-of-thought(title,abstract);reasoning(title);CoT(abstract);分类 cs.AI

Comments Minor modifications in sections: abstract, introduction, background problem formulation, and conclusion. (Typos and Clarifications)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05599 2025-06-10 cs.CV cs.CL 88%

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Yi Peng, Peiyu Wang, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, Rongxian Zhuang, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Skywork AI(Skywork人工智能) Kunlun Inc.(昆仑公司)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24181 2025-06-02 cs.AI 88%

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought

Guanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li, Hao Yu, Shuting Dong, Tao Ren, Ming Tang, Chun Yuan

机构 * SIGS, Tsinghua University(清华大学信息科学与技术学院) Southern University of Science and Technology(南方科技大学) Guangdong Laboratory of AI and Digital Economy (SZ)(广东人工智能与数字经济实验室) The Hong Kong University of Science and Technology(香港科技大学) Peking University(北京大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title);CoT(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21327 2025-05-28 cs.AI cs.CV 88%

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

Jiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu, Renrui Zhang, Kaituo Feng, Chaoyou Fu, Tao Chen, Lei Bai, Bo Zhang, Xiangyu Yue

机构 * Fudan University(复旦大学) MMLab, The Chinese University of Hong Kong(中大香港人工智能实验室) Shanghai AI Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学)

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08819 2024-12-13 cs.LG 88%

HARP: A challenging human-annotated math reasoning benchmark

Albert S. Yue, Lovish Madaan, Ted Moskovitz, DJ Strouse, Aaditya K. Singh

专题命中 推理评测 :reasoning(title,abstract);math reasoning(title,abstract);分类 cs.LG

Comments 28 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00559 2024-05-22 cs.CL 88%

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL

Comments Accepted to ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16049 2024-03-26 cs.CL 88%

MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, Greg Durrett

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL

Journal ref ICLR 2024 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14536 2023-10-24 cs.CL 88%

MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems

Jakub Macina, Nico Daheim, Sankalan Pal Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan

专题命中 推理评测 :reasoning(title,abstract);math reasoning(title,abstract);分类 cs.CL

Comments Jakub Macina, Nico Daheim, and Sankalan Pal Chowdhury contributed equally to this work. Accepted at EMNLP2023 Findings. Code and dataset available: https://github.com/eth-nlped/mathdial

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10244 2022-03-22 cs.CL 88%

ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, Enamul Hoque

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.CL

Comments Accepted by ACL 2022 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.08124 2020-07-17 cs.CL 88%

LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, Yue Zhang

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.CL

Comments Accepted by IJCAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22816 2026-04-14 cs.CL cs.AI cs.LG 88%

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness

衡量和治愈推理刚性:从装饰性推理链到真正的忠实

Abhinaba Basu, Pavan Chakraborty

机构 * Indian Institute of Information Technology Allahabad (IIITA)(印度阿拉哈巴德信息技术学院) National Institute of Electronics and Information Technology (NIELIT)(国家电子与信息技术研究院)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出SLRC指标和LC-CoSR方法,通过Lyapunov稳定性保证减少推理刚性,发现高SLRC模型易受阿谀影响,并提出RIS评分预测错误检测。

Comments Includes SLRC metric with formal guarantees (Theorem 1), LC-CoSR training intervention, Reasoning Integrity Score, and mechanistic analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11140 2026-08-05 cs.SE cs.AI cs.CL cs.HC 版本更新 88%

Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization

通过多路径推理和反馈驱动优化实现自动化可视化代码合成

Wonduk Seo, Daye Kang, Hyunjin An, Taehan Kim, Soohyuk Cho, Seungyong Lee, Minhyeong Yu, Jian Park, Yi Bu, Seunghyun Lee

机构 * AI Research, Enhans, Seoul, South Korea Innovation \& Technology, KAIST, Daejeon, South Korea Department of Computer Science, University of California, Berkeley, CA, United States Department of Electrical

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 VisPath通过多路径推理和反馈驱动优化,提升自动化可视化代码生成的可靠性与准确性。

Comments Accepted by International Conference on Pattern Recognization (ICPR 2026)

Journal ref Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science, vol 16812. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25182 2026-07-29 cs.CL cs.AI cs.IR 新提交 88%

TabRank: Chain-of-Thought Distillation for Table Re-Rankers

TabRank:表格重排器的思维链蒸馏

Adarsh Singh, Kushal Raj Bhandari, Jianxi Gao, Soham Dan, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) Rensselaer Polytechnic Institute(伦斯勒理工学院) Scale AI

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 研究针对表格检索训练推理重排器的问题,提出TabRank框架,通过构建数据集并探索两种训练紧凑推理模型的变体,经压力测试,该方法在多表格检索数据集上显著提升性能,有效推广到多表格推理。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17532 2026-06-02 cs.CL cs.LG 88%

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

OncoReason: 在大语言模型中构建临床推理以实现稳健且可解释的生存预测

Raghu Vamshi Hemadri, Geetha Krishna Guruju, Kristi Topollai, Anna Ewa Choromanska

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 提出统一多任务学习框架,通过监督微调、思维链提示和强化学习三种对齐策略,使自回归大语言模型在MSK-CHORD数据集上联合进行二元生存分类、连续生存时间回归和自然语言理由生成,实现可解释的生存预测。

Comments This manuscript is withdrawn to allow careful review and correction of bibliographic issues identified after submission, including references that could not be adequately verified. These matters should be resolved before further circulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10511 2026-05-29 cs.AI cs.CL 88%

Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation

快思考,错思考:直觉性调节LLM在政策评估中的反事实推理

Yanjie He

机构 * Independent Researcher(独立研究者)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 本研究构建了一个基于经济学和社会科学实证案例的基准,通过8000次实验评估大型语言模型在政策评估中的反事实推理,发现链式思维提示在反直觉案例中效果显著减弱,且直觉性是主导因素,表明模型存在知识-推理分离。

Comments 10 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13988 2026-05-13 cs.AI cs.LG 88%

Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning

忠实还是只是合理?评估闭源LLM在医学推理中的忠实性

Halimat Afolabi, Zainab Afolabi, Elizabeth Friel, Jude Roberts, Antonio Ji-Xu, Lloyd Chen, Egheosa Ogbomo, Emiliomo Imevbore, Phil Eneje, Wissal El Ouahidi, Aaron Sohal, Alisa Kennan, Shreya Srivastava, Anirudh Vairavan, Laura Napitu, Katie McClure

机构 * Stratified Precision Harvard Medical School(哈佛医学院) Imperial College London(帝国理工学院伦敦分校) National Health Service(国家健康服务系统) Ipsen France(Ipsen法国) University College London(伦敦大学学院)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 本文通过系统性黑盒评估,揭示闭源LLM在医学推理中的忠实性问题,发现链式推理步骤未必驱动预测,模型易受外部提示影响,强调忠实性比准确性更重要。

Journal ref Proceedings of Machine Learning Research, Vol. 297, pp. 1562-1591, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09008 2026-05-12 cs.LG cs.CL 88%

Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models

相对动能效用用于推理感知的结构剪枝在大语言模型中

Tianhao Qian

机构 * School of Mathematics, Southeast University(东南大学数学学院)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 本文提出相对动能效用框架,通过连续动能积分提升结构剪枝效果,实验证明在高稀疏度下提升模型性能,尤其在GSM8K数据集上表现优异。

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18806 2026-02-24 cs.CL cs.AI 88%

Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models

Think$^{2}$: 大语言模型中的 grounded 元认知推理

Abraham Paul Elenjical, Vivek Hruday Kavuri, Vasudeva Varma

机构 * IIIT Hyderabad(IIIT海得拉尔)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);planning(abstract);self-correction(abstract)

AI总结 Think$^{2}$通过引入基于心理理论的元认知框架,提升大语言模型的自我诊断与纠正能力,显著提高推理准确性和可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16975 2026-06-16 cs.SD eess.AS 版本更新 88%

Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs

基于多模态大语言模型的链式思维差异-共性推理的可解释音频编辑评估

Yuhang Jia, Xu Zhang, Yang Chen, Hui Wang, Enzhi Wang, Yong Qin

机构 * College of Computer Science, Nankai University, Tianjin, China(南开大学计算机科学学院,天津,中国) Academy for Advanced Interdisciplinary Studies, Nankai University, Tianjin, China(南开大学先进跨学科研究学院,天津,中国)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract)

AI总结 提出首个基于自然语言的音频编辑自动评估框架,利用Qwen2-Audio和链式思维推理策略,实现可解释且与人类判断高度一致的评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15242 2026-02-19 cs.CV 88%

Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities

可信且公平的SkinGPT-R1:用于在不同种族中普及皮肤病推理

Yuhao Shen, Zhangtianyi Chen, Yuanhao He, Yan Xu, Shuping Zhang, Liyuan Sun, Zijian Wang, Yinghao Zhu, Yuyuan Yang, Jiahe Qian, Ziwen Wang, Xinyuan Zhang, Wenbin Liu, Zongyuan Ge, Tao Lu, Siyuan Yan, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(数据科学学院,香港中文大学(深圳)) Faculty of Information Technology, Monash University(信息技术学院,莫纳什大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究所,天津中医研究院附属医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院)

专题命中 推理评测 :reasoning(title,abstract);logical reasoning(title);chain-of-thought(abstract)

AI总结 SkinGPT-R1通过公平性意识的专家混合架构实现可解释且公平的皮肤病诊断,提升不同种族间的诊断准确性与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02779 2025-11-05 cs.CV 88%

When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought

Yiyang Zhou, Haoqin Tu, Zijun Wang, Zeyu Wang, Niklas Muennighoff, Fan Nie, Yejin Choi, James Zou, Chaorui Deng, Shen Yan, Haoqi Fan, Cihang Xie, Huaxiu Yao, Qinghao Ye

机构 * ByteDance Seed(字节跳动种子) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) UC Santa Cruz(圣克拉拉大学) Stanford(斯坦福大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title);CoT(abstract)

Comments 28 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18700 2025-09-24 cs.SD eess.AS 88%

Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning

Chih-Cheng Chang, Bo-Yu Chen, Lu-Rong Chen, Li Su

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15651 2025-06-19 cs.LG cs.AI cs.CL 88%

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Tevin Wang, Chenyan Xiong

机构 * School of Computer Science(计算机科学学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08817 2025-06-13 cs.CV 88%

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Shuyi Zhang, Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Pengwei Wang, Zhongyuan Wang, Hongxuan Ma, Shanghang Zhang

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artifcial Intelligence, UCAS(中国科学技术大学人工智能学院) Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院) Shenzhen International GraduateSchool,Tsinghua University(深圳国际研究生院,清华大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院)

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16996 2024-03-26 cs.CV cs.RO 88%

DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving

Tianqi Wang, Enze Xie, Ruihang Chu, Zhenguo Li, Ping Luo

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24328 2026-04-20 cs.CL cs.AI 87%

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

超越多项选择题:一个包含方言变体的开放性阿拉伯文化问答基准测试

Hunzalah Hassan Bhatti, Firoj Alam

机构 * Qatar Computing Research Institute, Qatar(卡塔尔计算研究所)

专题命中 推理评测 :CoT(summary_cn,abstract);chain-of-thought(abstract,comments);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种方法,将现代标准阿拉伯语的多项选择题翻译成英文和阿拉伯方言,并转换为开放性问题,评估不同模型在不同设置下的表现,发现模型在阿拉伯方言上表现欠佳,CoT方法提高了正确性但指标混杂。

Comments Cultural Knowledge, Everyday Knowledge, Open-Ended Question, Chain-of-Thought, Large Language Models, Native, Multilingual, Language Diversity

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11470 2026-08-12 cs.CL 版本更新 87%

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

LLM推理的周期表:推理范式、方法与失败模式的结构化综述

Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar, Rishitej Reddy Vyalla, Erik Cambria, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah

机构 * Singapore Institute of Technology(新加坡理工大学) Nvidia AI Center (SNAIC)(英伟达人工智能中心(SNAIC)) MIDAS Lab, IIIT Delhi(IIIT德里MIDAS实验室) MIDAS Lab, IIT Mandi(IIT曼迪MIDAS实验室) Owl Autonomous Imaging, Inc.(Owl自主成像公司) College of Computing & Data Science, NTU Singapore(新加坡南洋理工大学计算与数据科学学院) NVIDIA AI Technology Centre, Singapore(英伟达新加坡人工智能技术中心) Department of Computer Science and Engineering, IIT Kanpur(IIT坎普尔计算机科学与工程系)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);verifier(abstract)

AI总结 本文系统综述了300多篇论文,提出LLM推理研究的结构化分类法,涵盖多种推理范式,分析方法论趋势,并总结常见限制与失败模式,旨在为开发更鲁棒、可解释和可泛化的推理系统提供参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26952 2026-07-30 cs.CL 新提交 87%

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

信用卡、困惑、计算与后果:我们能从语言模型推理中发现什么?

Arnav Hiray, Agam Shah, Caleb Lu, Meghaj Tarte, Harsit Mittal, Sudheer Chava

机构 * College of Computing(计算机学院) Scheller College of Business(舍勒商学院) Georgia Institute of Technology(佐治亚理工学院)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL

AI总结 该研究构建了首个基于真实信用卡协议的数值推理金融素养基准CreditCardQA,评估发现程序思维(PoT)提示可提升模型推理性能,错误多源于金融规则误用等,边缘案例易影响财务脆弱群体。

Comments Accepted at CoLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23458 2026-07-28 cs.CL 新提交 87%

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

思维链不忠实的两种模式:模型错误时行为检测失效

Suramya R. Angdembay, Dikshant Aryal, Nick Rahimi

机构 * The University of Southern Mississippi(南密西西比大学)

专题命中 推理评测 :chain-of-thought(title,abstract);CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL

AI总结 研究思维链不忠实检测,发现答案正确性影响检测效果,分正确与错误答案两种模式。仅答案不正确性表现优,跨模式无共享正向对齐方向,指示轨迹难转移,还解决了基准标签语义不匹配问题。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏