Think When You Need: Self-Adaptive Chain-of-Thought Learning
机构 * Xiaohongshu Inc(小红书公司)
专题命中 推理评测 :chain-of-thought(title);reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Under review
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
机构 * Xiaohongshu Inc(小红书公司)
专题命中 推理评测 :chain-of-thought(title);reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Under review
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Kupas Technology Limited Company(上海库帕斯科技有限公司)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)
机构 * College of Computing \& Data Science, Nanyang Technological University, Singapore
专题命中 推理评测 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract)
专题命中 推理评测 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)
Comments 11 pages, 5 figures
专题命中 推理评测 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments CVPR 2025 (Conference on Computer Vision and Pattern Recognition) Project page at https://jmhb0.github.io/microvqa Benchmark at https://huggingface.co/datasets/jmhb/microvqa
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);test-time compute(abstract)
专题命中 推理评测 :reasoning(title,abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2024
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Under Review
专题命中 推理评测 :reasoning(title,abstract);planning(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 推理评测 :reasoning(title,abstract);math reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint. Nico Daheim and Jakub Macina contributed equally. Code and dataset can be found under: https://github.com/eth-lre/verify-then-generate
专题命中 推理评测 :chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments TMLR accepted paper camera-ready version. First two authors contributed equally. 8 pages main, 13 pages appendix
专题命中 推理评测 :reasoning(title,abstract);self-correction(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL 2024 Findings
专题命中 推理评测 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2024 spotlight; 38 pages; code is at aka.ms/dyval
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ICLR 2024
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2023; updated with CLadder dataset v1.5
专题命中 推理评测 :planning(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
Comments arXiv admin note: text overlap with arXiv:2206.10498
专题命中 推理评测 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to NeurIPS 2022. 22 pages, 17 figures, 9 tables. Project: https://scienceqa.github.io
专题命中 推理评测 :chain-of-thought(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
PolyMATH:一个具有挑战性的多模态数学推理基准
机构 * Arizona State University(亚利桑那州立大学) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 PolyMATH基准通过5000张高质量图像评估多模态大语言模型的推理能力,揭示其在空间关系和抽象推理上的不足,指出模型无法真正理解视觉信息,存在逻辑错误风险。
Comments Accepted in Neural Information Processing Systems (NeurIPS 2025) Workshop: Foundations of Reasoning in Language Models
用于评估和加固LLM系统指令对抗编码攻击的自动化框架
机构 * Keysight Technologies
专题命中 推理评测 :chain-of-thought(summary_cn,abstract);reasoning(abstract);分类 cs.AI
AI总结 本文提出自动化框架评估LLM系统指令在对抗编码攻击时的保密性,通过四个模型和46条指令测试发现结构化序列化攻击成功率高,提出基于Chain-of-Thought的缓解策略。
Comments An earlier version of this manuscript will appear in the proceedings of IEEE Cyber-AI 2026 Conference. Project source code is available at https://github.com/Keysight/LLM-EncodeGuard
CRAFT:基于大语言模型的迭代优化用于临床叙事的时间推理
机构 * Stevens Institute of Technology(史蒂文斯理工学院) ; Genesis Research Group(创世纪研究集团) ; New Jersey Institute of Technology(新泽西理工学院)
专题命中 推理评测 :reasoning(title,abstract);verifier(abstract);分类 cs.CL、cs.AI
AI总结 针对临床叙事中时间锚点稀疏导致的症状轨迹重建难题,提出基于大语言模型的CRAFT框架,通过生成器与约束验证器迭代优化,在MedTempo基准上提升了时间排序准确率。
V-FiLLM:经过验证的金融大语言模型推理基准
机构 * ETH Zürich(苏黎世联邦理工学院) ; Aisot Technologies Ltd(艾索特科技有限公司)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG
AI总结 本文提出V-FiLLM框架生成高可靠性金融推理基准,发现开源模型在金融推理深度增加及数值扰动下准确率显著下降,轻量级LoRA微调可提升模型在相关任务的表现。
Comments 10 pages, 6 tables, 2 figures, under review
从泄露思维到私有推理:控制LRM对自己说的话
机构 * Mohamed bin Zayed University of Artificial Intelligence, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 针对大型推理模型(LRM)推理过程中隐私泄露问题,提出通过指令跟随(IF)训练和分阶段解码策略(Staged Decoding)增强隐私保护,在IF和隐私基准上分别提升高达20.9和51.9个百分点。
推理大语言模型中的测试时缩放:推理机制、评估与可复现性
专题命中 推理评测 :reasoning(title,abstract);verifier(abstract);分类 cs.AI、cs.LG
AI总结 该研究针对推理大语言模型的测试时缩放,明确三种推理机制,提出系统评估原则与可复现性要求,应用于多类基准并发布超20亿条推理轨迹。
量化语言模型推理中的无声失败:基于分类法的空洞收敛和失败模式转移分析
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG
AI总结 研究量化大语言模型推理中的无声失败,用六类失败分类法对五个模型在三种精度和四个基准下的思维链输出分类,发现空洞收敛等问题有精度和基准依赖性,且无法从文本特征可靠检测,揭示了标准评估管道未捕捉到的失败模式。
IsoSci: 用于评估LLM中推理与知识检索的同构跨域科学问题基准
机构 * Texas A&M University(德克萨斯农工大学) ; Hamad Bin Khalifa University(哈马德·本·哈利法大学)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 提出IsoSci基准,通过逻辑结构相同但领域知识不同的科学问题对,分离推理能力与领域知识检索,发现91.3%的推理增益依赖于知识而非结构,挑战了思维链推理提升科学问题解决的假设。