arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10451 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10451 篇

2507.21432 2025-10-08 cs.CL cs.AI 62%

Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour

Tareq Alsaleh, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Canada(交通创新实验室(LiTrans)、多伦多 Metropolitan 大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05151 2025-10-08 cs.CL cs.LG 62%

Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA

Prudence Djagba, Abdelkader Y. Saley

机构 * Lyman Briggs College, Michigan State University(密歇根州立大学Lyman Briggs学院) Department of Finance, Michigan State University(密歇根州立大学金融系) African Institute for Mathematical Sciences, Rwanda(刚果(金)数学科学研究所)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05135 2025-10-08 cs.CL cs.LG 62%

Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment

Vanya Bannihatti Kumar, Divyanshu Goyal, Akhil Eppa, Neel Bhandari

机构 * Adobe Inc.(Adobe公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22385 2025-10-08 cs.CV cs.AI cs.CL 62%

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

Yue Zhang, Jilei Sun, Yunhui Guo, Vibhav Gogate

机构 * Department of Computer Science(计算机科学系) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05040 2025-10-07 cs.LG cs.AI 62%

Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts

Jihoon Lee, Hoyeon Moon, Kevin Zhai, Arun Kumar Chithanar, Anit Kumar Sahu, Soummya Kar, Chul Lee, Souradip Chakraborty, Amrit Singh Bedi

机构 * Yonsei University(延世大学) Oracle(Oracle公司) CMU(卡内基梅隆大学) UMD(马里兰大学) UCF(佛罗里达大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04980 2025-10-07 cs.AI cs.CL 62%

LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game

Fangzhou Liang, Tianshi Zheng, Chunkit Chan, Yauwai Yim, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Wordplay

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04506 2025-10-07 cs.CL cs.AI cs.IR 62%

GRACE: Generative Representation Learning via Contrastive Policy Optimization

Jiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong, Pengcheng Jiang, Bowen Jin, Peiran Li, Weijia Shi, Jiawei Han

机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) Australian National University(澳大利亚国立大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 23 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04374 2025-10-07 cs.LG cs.AI cs.CY 62%

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simón Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, Jerry Tworek

机构 * OpenAI

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04320 2025-10-07 cs.CL cs.LG 62%

Read the Scene, Not the Script: Outcome-Aware Safety for LLMs

Rui Wu, Yihao Quan, Zeru Shi, Zhenting Wang, Yanshu Li, Ruixiang Tang

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04032 2025-10-07 cs.CL cs.AI 62%

Small Language Models for Emergency Departments Decision Support: A Benchmark Study

Zirui Wang, Jiajun Wu, Braden Teitge, Jessalyn Holodinsky, Steve Drew

机构 * Department of Electrical and Software Engineering, University of Calgary, Calgary, AB, Canada(电气与软件工程系,卡尔加里大学) Department of Emergency Medicine, University of Calgary, Calgary, AB, Canada(急诊医学系,卡尔加里大学) Rockview General Hospital, Calgary, AB, Canada(罗克维尔医院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to 2025 IEEE International Conference on Autonomous and Trusted Computing (ATC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22830 2025-10-07 cs.CL cs.AI 62%

What Has Been Lost with Synthetic Evaluation?

Alexander Gill, Abhilasha Ravichander, Ana Marasović

机构 * University of Utah(犹他大学) University of Washington(华盛顿大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments v3: Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01159 2025-10-07 cs.LG cs.AI 62%

AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science

Chenyue Li, Wen Deng, Mengqian Lu, Binhang Yuan

机构 * HKUST(香港科技大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 37 pages, 4 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18917 2025-10-07 cs.LG cs.AI 62%

Behavior Injection: Preparing Language Models for Reinforcement Learning

Zhepeng Cen, Yihang Yao, William Han, Zuxin Liu, Ding Zhao

机构 * Carnegie Mellon University(卡内基梅隆大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02830 2025-10-06 cs.CL cs.AI 62%

Evaluating Large Language Models for IUCN Red List Species Information

Shinya Uryu

机构 * Center for Design-Oriented AI Education and Research(设计导向人工智能教育与研究中心) Tokushima University(德岛大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02549 2025-10-06 cs.CL cs.AI 62%

Knowledge-Graph Based RAG System Evaluation Framework

Sicheng Dong, Vahid Zolfaghari, Nenad Petrovic, Alois Knoll

机构 * Technical University of Munich, Robotics, Artificial Intelligence and Embedded Systems(慕尼黑技术大学,机器人,人工智能与嵌入式系统)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02326 2025-10-06 cs.CL cs.AI 62%

Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval

Vivek Bhavsar, Joseph Ereifej, Aravanan Gurusami

机构 * CTO Office, Coherent Corporation(Coherent Corporation 技术总监办公室)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01227 2025-10-06 cs.CL cs.LG math.HO 62%

EEFSUVA: A New Mathematical Olympiad Benchmark

Nicole N Khatibi, Daniil A. Radamovich, Michael P. Brenner

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

Comments 16 Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14052 2025-10-06 cs.IR cs.AI cs.CL 62%

FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering

Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim, Minjae Kim, Juneha Hwang, Jaeseon Ha, Hojun Choi, Suyeol Yun, Yongjin Kim, Yongjae Lee

机构 * University of Florida(佛罗里达大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01232 2025-10-03 cs.CL cs.AI 62%

Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks

Dongjun Kim, Gyuho Shim, Yongchan Chun, Minhyuk Kim, Chanjun Park, Heuiseok Lim

机构 * Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系) School of Software, Soongsil University(顺天大学软件学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 5 figures. Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19430 2025-10-03 cs.CL cs.AI 62%

Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation

Keane Ong, Rui Mao, Deeksha Varshney, Paul Pu Liang, Erik Cambria, Gianmarco Mengaldo

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Published at Empirical Methods in Natural Language Processing 2025 (Main Conference) (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00428 2025-10-02 cs.LG cs.AI 62%

Automated Structured Radiology Report Generation with Rich Clinical Context

Seongjae Kang, Dong Bok Lee, Juho Jung, Dongseop Kim, Won Hwa Kim, Sunghoon Joo

机构 * VUNO Inc.(VUNO公司) KAIST(韩国科学技术院) POSTECH(POSTECH大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 34 pages, 30 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26417 2025-10-01 cs.AI cs.LG 62%

OntoAligner Meets Knowledge Graph Embedding Aligners

Hamed Babaei Giglou, Jennifer D'Souza, Sören Auer, Mahsa Sanaei

机构 * TIB -- Leibniz Information Centre for Science and Technology(莱比锡信息科学与技术研究中心) University of Tabriz(塔布里兹大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 10 pages of main content, 3 page references, 3 figures. Accepted to Ontology Matching Workshop at ISWC

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25813 2025-10-01 cs.CL cs.LG 62%

RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models

Dragos-Dumitru Ghinea, Adela-Nicoleta Corbeanu, Adrian-Marius Dumitran

机构 * University of Bucharest(布加勒斯特大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25238 2025-10-01 cs.LG cs.AI 62%

PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases

Sri Vatsa Vuddanti, Aarav Shah, Satwik Kumar Chittiprolu, Tony Song, Sunishchal Dev, Kevin Zhu, Maheep Chaudhary

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24297 2025-10-01 cs.CL cs.AI 62%

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen, Yalun Wu, Yingji Liang, Yijin Guo, Farong Wen, Wenzhe Li, Xuezhi Zhao, Qi Jia, Guangtao Zhai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14017 2025-10-01 cs.CL cs.LG 62%

Efficient Temporal Tokenization for Mobility Prediction with Large Language Models

Haoyu He, Haozheng Luo, Yan Chen, Qi R. Wang

机构 * Northeastern University, Boston, MA(东北大学) Northwestern University, Evanston, IL(西北大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

Journal ref Proceedings of the 3rd Workshop on Efficient Systems for Foundation Models (ES-FoMo III) at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21500 2025-10-01 cs.CV cs.AI cs.CL 62%

ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models

Dingming Li, Hongxing Li, Zixuan Wang, Yuchen Yan, Hang Zhang, Siqi Chen, Guiyang Hou, Shengpei Jiang, Wenqi Zhang, Yongliang Shen, Weiming Lu, Yueting Zhuang

机构 * Zhejiang University(浙江大学) University of Electronic Science and Technology of China(电子科技大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Project: https://zju-real.github.io/ViewSpatial-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04690 2025-09-30 cs.LG cs.AI 62%

Towards Better Generalization via Distributional Input Projection Network

Yifan Hao, Yanxin Lu, Hanning Zhang, Xinwei Shen, Tong Zhang

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03930 2025-09-30 cs.SE cs.AI cs.CL 62%

VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation

Yuansheng Ni, Ping Nie, Kai Zou, Xiang Yue, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :self-correction(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23267 2025-09-30 cs.CV cs.AI cs.LG 62%

Learning Regional Monsoon Patterns with a Multimodal Attention U-Net

Swaib Ilias Mazumder, Manish Kumar, Aparajita Khan

机构 * 1Computer Science \& Engineering, Indian Institute of Technology Roorkee, India 2Computer Science \& Engineering, Indian Institute of Technology Ropar, India 3 Computer Science \& Engineering, Indian Institute of Technology (BHU) Varanasi, India

专题命中 推理评测 :planning(abstract);分类 cs.AI、cs.LG

Comments Accepted in Geospatial AI and Applications with Foundation Models (GAIA) 2025, INSAIT and ELLIS Unit Sofia, Bulgaria

详情

展开后加载摘要…

URL PDF HTML 收藏