arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-04-21 至 2026-04-21 共收录 91 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 91 篇

2604.10693 2026-04-21 cs.AI 90%

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

FACT-E:基于因果的可信推理链评估

Yuxi Sun, Aoqi Zuo, Haotian Xie, Wei Gao, Mingming Gong, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) The University of Melbourne(墨尔本大学) Singapore Management University(新加坡管理大学)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract,abstract_cn);分类 cs.AI

AI总结 FACT-E通过因果方法评估推理链的可信度,结合内部一致性与推理-答案一致性,提升推理轨迹选择质量并增强噪声下的鲁棒性。

Comments Accepted to Association for Computational Linguistics Findings (ACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11113 2026-04-21 cs.CV cs.AI cs.LG 90%

VIDEOP2R: Video Understanding from Perception to Reasoning

VIDEOP2R:从感知到推理的视频理解

Yifan Jiang, Yueying Wang, Rui Zhao, Toufiq Parag, Zhimin Chen, Zhenyu Liao, Jayakrishnan Unnikrishnan

机构 * USC(USC大学) Amazon(亚马逊) Keystone AI

专题命中 推理评测 :CoT(summary_cn,abstract);reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 VIDEOP2R提出一种过程感知的视频RFT框架,通过三步流程生成高质量CoT数据集,并引入PA-GRPO算法提升视频推理性能,实现在七个基准中的六个达到SotA水平。

Comments CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11281 2026-04-21 cs.CL cs.AI cs.CY 89%

ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection

ToxiFrench:通过CoT微调提升法语毒性检测的语言模型基准测试

Axel Delaval, Shujian Yang, Haicheng Wang, Han Qiu, Jialiang Lu

机构 * École Polytechnique(巴黎高等理工学院) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学)

专题命中 推理评测 :CoT(title,title_cn);chain-of-thought(abstract,comments);分类 cs.CL、cs.AI

AI总结 本文提出ToxiFrench数据集,通过半自动化标注流程构建,发现小语言模型在毒性检测任务中表现更优,并提出动态加权损失策略提升模型忠实度,Qwen3-4B模型在基准测试中取得最佳性能。

Comments 22 pages, 5 figures, 11 tables. This paper introduces TOXIFRENCH, a benchmark of 53,622 comments for French toxicity detection. It proposes a Chain-of-Thought fine-tuning method with a dynamic weighted loss. The fine-tuned 4B model (Qwen3-4B) achieves state-of-the-art performance, outperforming larger models like GPT-4o and DeepSeek-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12782 2026-04-21 cs.AI 88%

HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds

HeroBench:一个用于虚拟世界中长horizon规划和结构化推理的基准测试

Petr Anokhin, Roman Khalikov, Stefan Rebrikov, Viktor Volkov, Artyom Sorokin, Vincent Bissonnette

机构 * Artyom Sorokin(AXXX) Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学) Higher School of Economics(高等经济学院) Kurchatov Institute(库尔斯克研究所) Independent Researcher(独立研究者)

专题命中 推理评测 :reasoning(title,abstract);planning(title,abstract);分类 cs.AI

AI总结 本文提出HeroBench,用于评估在复杂RPG-inspired虚拟世界中长horizon分层规划和结构化推理的能力,通过模拟评估可执行计划,揭示了现有大型语言模型在长horizon自主规划中的性能差异和挑战。

Comments Code is available at https://github.com/stefanrer/HeroBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17295 2026-04-21 cs.AI 87%

LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics

LLaTiSA:从视觉感知到语义的难度分层时间序列推理

Yueyang Ding, HaoPeng Zhang, Rui Dai, Yi Wang, Tianyu Zong, Kaikui Liu, Xiangxiang Chu

机构 * Alibaba Group(阿里巴巴集团) Amap(高德)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出LLaTiSA,通过四层分类体系提升时间序列推理能力,结合可视化模式与精确校准的数值表格增强视觉语言模型的时序感知,实现跨多样任务和现实场景的鲁棒泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07993 2026-04-21 cs.AI 87%

SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models

SkipKV: 为大推理模型高效推理选择性跳过KV生成与存储

Jiayi Tian, Seyedarmin Azizi, Yequan Zhao, Erfan Baghaei Potraghloo, Sean McPherson, Sharath Nittur Sridhar, Zhengyang Wang, Zheng Zhang, Massoud Pedram, Souvik Kundu

机构 * Anonymous Institution, Anonymous City, Anonymous Region, Anonymous Country(匿名机构,匿名城市,匿名地区,匿名国家)

专题命中 推理评测 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出SkipKV方法,通过选择性跳过KV生成与存储,减少缓存大小,提升推理效率。方法引入句子评分指标,动态调整引导向量,实现高效推理,实验显示准确率提升26.7%,生成长度缩短1.6倍,吞吐量提升1.7倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17794 2026-04-21 cs.CL cs.AI 86%

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

通过测试时间扩展弥合越南语中的推理差距

Bui The Trung, Do Minh Duc, Nguyen Van Vinh, Bui Nguyen Quoc Trinh

机构 * Computer Science and Engineering(计算机科学与工程) Information Science Faculty(信息科学系) Faculty of Information(信息学院) Faculty of Advanced Technologies and Engineering(先进技术和工程学院) VNU Vietnam(越南VNU) Japan University(日本大学) JAIST VNU University of Engineering and Technology(VNU工程大学) VNU Vietnam Japan University(越南日本大学) Hanoi, Vietnam(越南河内) Ishikawa, Japan(日本石川)

专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了在越南小学数学中通过测试时间扩展策略提升小语言模型推理能力,提出Vi-S1K和Vi-Elementary-Bench基准,发现SFT显著提升解释质量,证明简化测试时间扩展优于复杂代理流程。

Comments FJICAI conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17282 2026-04-21 cs.CL 83%

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

MedPRMBench: 一种面向医疗推理过程奖励模型的细粒度基准

Lingyan Wu, Xiang Zheng, Weiqi Zhai, Wei Wang, Xuan Ren, Zifan Zhang, Hu Wei, Bing Zhao

机构 * Alibaba Group(阿里巴巴集团)

专题命中 推理评测 :reasoning(title,abstract);verifier(abstract);分类 cs.CL

AI总结 本文提出MedPRMBench,首个医疗领域过程奖励模型基准,通过三阶段流程生成高质量数据,涵盖14种细粒度错误类型,评估模型在医疗推理中的误差检测能力,提升下游任务准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21278 2026-04-21 cs.CV cs.AI cs.CL cs.LG 82%

GeoRC: A Benchmark for Geolocation Reasoning Chains

GeoRC:地理定位推理链基准测试

Mohit Talreja, Joshua Diao, Jim Thannikary James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出GeoRC基准,通过Champion级GeoGuessr专家生成的800条真实推理链,评估VLM生成推理链的准确性,发现大模型在生成可审计推理链上仍逊于人类专家,而小型模型表现更差。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07562 2026-04-21 cs.CL cs.AI cs.CY cs.LG 82%

Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs

基于推理的无监督文本聚类细化方法

Tunazzina Islam

机构 * Department of Computer Science(计算机科学系)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出利用大语言模型作为语义评判者,对无监督聚类结果进行推理细化,通过三个阶段提升聚类的连贯性与可解释性,实验证明其在社交媒体数据中优于传统方法。

Comments Accepted to the Findings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16857 2026-04-21 cs.CV cs.RO 82%

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

BOP-ASK:面向视觉-语言模型的对象交互推理

Vineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich, Prashanth Krishnamurthy, Ramesh Karri, Stan Birchfield, Farshad Khorrami, Jonathan Tremblay

机构 * New York University(纽约大学) NVIDIA(英伟达)

专题命中 推理评测 :reasoning(title,abstract);planning(abstract)

AI总结 本文提出BOP-ASK数据集,用于训练和评估视觉-语言模型的对象交互推理能力,包含15万张图像和3300万对问题答案,涵盖六个任务,验证了模型在复杂环境中的空间推理能力。

Comments Accepted at CVPR 2026. Code, Datasets & Benchmark available at https://bop-ask.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03331 2026-04-21 cs.CV cs.AI cs.LG 81%

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

MMErroR:面向视觉-语言模型错误推理的基准测试

Yang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu, Mingxuan Huang, Jingchao Wang, Zhihong Zhu, Boyan Xu, Zhiqi Huang

机构 * Guangdong University of Technology(广东工业大学) Hong Kong Baptist University(香港 Baptist大学) Sun Yat-sen University(中山大学) Peking University(北京大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出MMErroR基准,通过1997个包含单一推理错误的样本,评估视觉-语言模型检测错误推理的能力,发现即使最佳模型也仅能正确分类66.65%的错误。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23808 2026-04-21 cs.LG cs.CL 81%

Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning

在RLVR中对LLM推理进行语义空间探索与利用

Fanding Huang, Guanbo Huang, Xiao Fan, Yi He, Xiao Liang, Xiao Chen, Qinting Jiang, Faisal Nadeem Khan, Jingyan Jiang, Zhi Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) University of California, Los Angeles(加州大学洛杉矶分校) Shenzhen Technology University(深圳技术大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出VERL方法,通过有效秩和其时间导数改进RLVR,实现语义空间中探索与利用的平衡,提升LLM推理性能。

Comments Accepted as an ACL 2026 Findings paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11137 2026-04-21 cs.AI cs.LG 81%

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

从答案到论证:通过托尔金引导的课程目标条件学习实现可信的临床诊断推理

Chen Zhan, Xiaoyu Tan, Gengchen Ma, Yu-Jie Xiong, Xiaoyan Jiang, Xihe Qiu

机构 * School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院) Tencent Youtu Lab(腾讯优图实验室)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出托尔金引导的课程目标条件学习框架,通过逐步训练LLM生成符合托尔金结构的诊断论证,提升临床诊断的透明性和可靠性。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05336 2026-04-21 cs.CL cs.AI 81%

WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives

WeatherArchive-Bench: 基于历史天气档案的检索增强推理基准测试

Yongan Yu, Xianda Du, Qingchen Hu, Jiahao Liang, Jingwei Ni, Dan Qiang, Kaiyu Huang, Grant McKenzie, Renee Sieber, Fengran Mo

机构 * McGill University(麦吉尔大学) University of Waterloo(滑铁卢大学) ETH Zurich(苏黎世联邦理工学院) Beijing Jiaotong University(北京交通大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出WeatherArchive-Bench,用于评估检索增强生成系统在处理历史天气档案中的能力,揭示密集检索器在历史术语上的不足及LLM对社会脆弱性与韧性概念的误判问题。

Comments accepted to the Resource Track of SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16913 2026-04-21 cs.AI cs.CL cs.CR cs.DC 81%

The Cognitive Penalty: Ablating System 1 and System 2 Reasoning in Edge-Native SLMs for Decentralized Consensus

认知惩罚:在边缘原生SLM中消融系统1和系统2推理以实现去中心化共识

Syed Muhammad Aqdas Rizvi

机构 * Independent Researcher(独立研究者) Lahore University of Management Sciences(拉合尔管理科学大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文研究了边缘原生SLM中系统1和系统2推理对去中心化共识的影响,发现系统1在对抗性环境中表现更优,而系统2导致计算准确率倒置和稳定性下降。

Comments Working paper. 14 pages, 3 figures, 6 tables. Code and dataset: https://github.com/smarizvi110/sentinel-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18234 2026-04-21 cs.IR cs.AI 79%

Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies

评估基于检索增强生成系统的多跳推理:LLM检索评估策略的比较

Lorenz Brehme, Thomas Ströhle, Ruth Breu

机构 * Universität Innsbruck(因斯布鲁克大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文通过对比三种LLM评估策略,提出上下文感知检索评估(CARE),证明其在提升RAG系统多跳推理能力方面效果显著,尤其在参数多和上下文长的模型中表现更优。

Comments 15 Pages, Accepted for publication at the SynIRgy Workshop, ECIR 2026 (48th European Conference on Information Retrieval)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18176 2026-04-21 cs.AI quant-ph 79%

QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning

量子QA:通过物理一致的数据集和验证意识强化学习增强科学推理

Songxin Qu, Tai-Ping Sun, Yun-Jie Wang, Huan-Yu Liu, Cheng Xue, Xiao-Fan Xu, Han Fang, Yang Yang, Yu-Chun Wu, Guo-Ping Guo, Zhao-Yun Chen

机构 * Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) School of Physics, University of Science and Technology of China(中国科学技术大学物理学院) School of Electronics and Information Engineering, Anhui University(安徽大学电子与信息工程学院) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出量子QA数据集和验证意识奖励模型,通过物理一致的数据集和强化学习提升科学推理能力,实验证明其在科学领域表现优异。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12320 2026-04-21 cs.CV cs.AI cs.MM 79%

EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports

EgoEsportsQA:一种用于电子竞技感知与推理的视角视频基准

Jianzhe Ma, Zhonghao Cao, Shangkui Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出EgoEsportsQA基准,通过六阶段流程收集1745对高质量问答对,评估视频大语言模型在虚拟环境中的感知与推理能力,揭示模型在战术推理和微操作方面的不足。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05543 2026-04-21 cs.CL cs.SD eess.AS 79%

Closing the Modality Reasoning Gap for Speech Large Language Models

弥合语音大语言模型的模态推理差距

Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, Zhizheng Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Microsoft Corporation(微软公司)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文提出TARS框架,通过不对称奖励设计对齐文本和语音条件轨迹,显著缩小模态推理差距并在7B级语音大语言模型中取得最佳性能。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04809 2026-04-21 cs.AI 79%

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning

SCALER:合成可扩展自适应学习环境用于推理

Caijun Xu, Changyi Xiao, Zhongyuan Peng, Xinrun Wang, Yixin Cao

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Singapore Management University(新加坡国立大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 SCALER通过自适应环境设计维持有效学习信号,解决RL训练中任务难度与模型能力不匹配及训练模式狭窄的问题,提升推理能力。

Comments 22 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09351 2026-04-21 cs.CL 79%

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

ReTraceQA:评估小语言模型在常识问答中的推理轨迹

Francesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia, Roberto Navigli

机构 * Sapienza NLP Group(萨皮恩扎大学自然语言处理小组) Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 ReTraceQA通过引入过程级评估,揭示小语言模型在常识推理中存在大量推理过程错误,但最终答案正确的情况,表明现有评价方法高估了模型能力。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17186 2026-04-21 cs.SE cs.AI cs.ET cs.HC cs.MA 79%

Persona-Based Requirements Engineering for Explainable Multi-Agent Educational Systems: A Scenario Simulator for Clinical Reasoning Training

基于角色的可解释多智能体教育系统需求工程:用于临床推理训练的场景模拟器

Weibing Zheng, Laurah Turner, Jess Kropczynski, Matthew Kelleher, Murat Ozer, Shane Halse

机构 * School of Information Technology University of Cincinnati, OH Cincinnati, USA(信息科技学院 奥地利辛辛那提大学) College of Medicine University of Cincinnati, OH Cincinnati, USA(医学院 奥地利辛辛那提大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出基于角色的可解释多智能体教育系统需求工程框架,通过临床推理训练系统展示如何利用角色和用户故事捕捉多方需求,提升系统可解释性和临床场景真实性。

Comments 7 pages, 2 figures, CSTE2026: https://cste.net/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11161 2026-04-21 cs.HC cs.CL 79%

Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning

Althea:面向事实核查和批判性推理的人机协作

Svetlana Churina, Kokil Jaidka, Anab Maulana Barik, Harshit Aneja, Cai Yang, Wynne Hsu, Mong Li Lee

机构 * Centre for Trusted Internet & Community (CTIC)(可信互联网与社区中心) National University of Singapore (NUS)(新加坡国立大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 Althea通过整合问题生成、证据检索和结构化推理,提升在线主张的核查效率与准确性,其在AVeriteC基准测试中达到宏F1值0.44,且在用户研究中显示指导性交互对短期准确性和长期改进均有显著影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16929 2026-04-21 cs.CL 79%

MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

MeasHalu:通过增强推理缓解大型语言模型中的科学测量幻觉

Ruijun Huang, Zhiqiao Kang, Yuxuan Zhu, Junxiong Li, Jiahao Zhao, Minghuan Tan, Feng Jiang, Min Yang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(深圳先进技术大学人工智能研究院)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文提出MeasHalu框架,通过增强推理和针对性优化缓解LLM的科学测量幻觉问题,改进测量提取的准确性。

Comments To appear in ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16862 2026-04-21 cs.LG 79%

Learning to Trade Like an Expert: Cognitive Fine-Tuning for Stable Financial Reasoning in Language Models

学习像专家一样交易:用于语言模型在金融推理中稳定性的认知微调

Yuchen Pan, Soung Chang Liew

机构 * Department of Information Engineering, The Chinese University of Hong Kong(信息工程系,香港中文大学)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.LG

AI总结 本文提出一种框架,通过 curated 多选题数据集和两阶段评估协议,训练语言模型在金融决策中具备稳定性和风险意识,优于开源基线并接近前沿模型表现。

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18361 2026-04-21 cs.CL 79%

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models

为训练多样化常识推理模型生成合成数据

Tianhui Zhang, Bei Peng, Danushka Bollegala

机构 * University of Liverpool(利物浦大学) University of Sheffield(谢菲尔德大学) Amazon(亚马逊)

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL

AI总结 本文提出两阶段方法生成首个合成数据集CommonSyn,提升生成多样性和质量,适用于不同规模的大语言模型。

Comments ACL2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17305 2026-04-21 cs.CE 78%

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications

BizCompass:评估大语言模型在商业知识与应用中的推理能力的基准测试

Jianing Hao, Yuhe Wu, Yuanjian Xu, Shichang Meng, Shuai Yuan, Wei Zeng, Zixuan Wang, Guang Zhang

专题命中 推理评测 :reasoning(title,abstract)

AI总结 BizCompass通过连接理论基础与实际商业知识,评估LLM在金融、经济等领域的推理能力,揭示不同模型在商业场景中的表现差异及影响因素。

Comments 40 pages, 6 figures, Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04334 2026-04-21 cs.CV 78%

GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models

GeoArena:在大视觉-语言模型中评估开放世界地理推理

Pengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon Li

机构 * Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 本文提出GeoArena框架,通过人偏好评估动态图像中的地理推理能力,评估17种前沿LVLM,分析模型行为与人类偏好可靠性。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02930 2026-04-21 cs.CV 78%

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos

TemporalVLM:长视频中的时间推理视频大语言模型

Fawad Javed Fateh, Umer Ahmed, Hamza Khan, M. Zeeshan Zia, Quoc-Huy Tran

机构 * Retrocausal, Inc.(Retrocausal公司)

专题命中 推理评测 :reasoning(title,abstract)

AI总结 本文提出TemporalVLM,一种用于长视频时间推理和细粒度理解的视频大语言模型,通过时间感知编码和BiLSTM模块实现全局特征聚合,并引入IndustryASM数据集进行评估。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏