arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10414 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10414 篇

2404.13065 2024-04-23 cs.CL cs.AI 81%

Intellecta Cognitiva: A Comprehensive Dataset for Advancing Academic Knowledge and Machine Reasoning

Ajmal PS, Ditto PS, Jithin VG

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19426 2024-04-19 cs.CL cs.LG 81%

ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning

Jingyuan Selena She, Christopher Potts, Samuel R. Bowman, Atticus Geiger

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04302 2024-04-09 cs.CL cs.AI 81%

CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering

Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawardena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-Orji, Ruvan Weerasinghe, Anne Liret, Bruno Fleisch

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Submitted to ICCBR'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09702 2024-04-09 cs.CL cs.AI 81%

Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?

Bangzheng Li, Ben Zhou, Fei Wang, Xingyu Fu, Dan Roth, Muhao Chen

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Work accepted by NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10797 2024-04-08 cs.CL cs.AI 81%

TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes

Bibek Upadhayay, Vahid Behzadan

专题命中 推理评测 :chain-of-thought(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02330 2024-04-01 cs.AI cs.CL 81%

Enhance Reasoning for Large Language Models in the Game Werewolf

Shuang Wu, Liwen Zhu, Tao Yang, Shiwei Xu, Qiang Fu, Yang Wei, Haobo Fu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02477 2024-04-01 cs.CL cs.AI 81%

Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02496 2024-03-29 cs.CL cs.AI 81%

Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification

Shan Chen, Yingya Li, Sheng Lu, Hoang Van, Hugo JWL Aerts, Guergana K. Savova, Danielle S. Bitterman

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 28 pages, 2 tables and 4 figures. Submitting for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07527 2024-03-26 cs.CL cs.AI 81%

BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability

Peter Clark, Bhavana Dalvi Mishra, Oyvind Tafjord

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Added note about how dataset sampling was performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14255 2024-03-22 cs.CL cs.LG 81%

ERD: A Framework for Improving LLM Reasoning for Cognitive Distortion Classification

Sehee Lim, Yejin Kim, Chi-Hyun Choi, Jy-yong Sohn, Byung-Hoon Kim

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11335 2024-03-21 cs.RO cs.AI cs.CV cs.LG 81%

Surfer: Progressive Reasoning with World Models for Robotic Manipulation

Pengzhen Ren, Kaidong Zhang, Hetao Zheng, Zixuan Li, Yuhang Wen, Fengda Zhu, Mas Ma, Xiaodan Liang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19450 2024-03-01 cs.AI cs.CL 81%

Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Saurabh Srivastava, Annarose M B, Anto P, Shashank Menon, Ajay Sukumar, Adwaith Samod T, Alan Philipose, Stevin Prince, Sooraj Thomas

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 37 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14840 2024-02-26 cs.CL cs.AI stat.AP 81%

RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning

Congyun Jin, Ming Zhang, Xiaowei Ma, Li Yujiao, Yingbo Wang, Yabo Jia, Yuliang Du, Tao Sun, Haowen Wang, Cong Fan, Jinjie Gu, Chenfei Chi, Xiangguo Lv, Fangzhou Li, Wei Xue, Yiran Huang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13028 2024-02-21 cs.CL cs.AI 81%

Heterogeneous Graph Reasoning for Fact Checking over Texts and Tables

Haisong Gong, Weizhi Xu, Shu wu, Qiang Liu, Liang Wang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by 38th Association for the Advancement of Artificial Intelligence, AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06805 2024-01-19 cs.CL cs.AI 81%

Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, Hongxia Yang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02132 2024-01-05 cs.CL cs.AI 81%

DCR-Consistency: Divide-Conquer-Reasoning for Consistency Evaluation and Improvement of Large Language Models

Wendi Cui, Jiaxin Zhang, Zhuohang Li, Lopez Damien, Kamalika Das, Bradley Malin, Sricharan Kumar

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17661 2024-01-01 cs.CL cs.AI cs.CV 81%

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models

Yuqing Wang, Yun Zhao

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Data and results are available at: https://github.com/EternityYW/Gemini-Commonsense-Evaluation/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09247 2023-12-25 cs.AI cs.LG 81%

Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks

Melanie Mitchell, Alessandro B. Palmarini, Arseny Moskvichev

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments Corrected Figure 3 (extra spaces were replaced by commas, which were lost in original formatting)

Journal ref Proceedings of the LLM-CP Workshop, AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08506 2023-12-19 cs.CV cs.AI cs.LG 81%

Does Visual Pretraining Help End-to-End Reasoning?

Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab, Cordelia Schmid

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15448 2023-12-06 cs.CL cs.AI cs.HC 81%

Understanding Social Reasoning in Language Models with Language Models

Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, Noah D. Goodman

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15930 2023-11-28 cs.CL cs.AI 81%

WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Youssef Benchekroun, Megi Dervishi, Mark Ibrahim, Jean-Baptiste Gaya, Xavier Martinet, Grégoire Mialon, Thomas Scialom, Emmanuel Dupoux, Dieuwke Hupkes, Pascal Vincent

专题命中 推理评测 :reasoning(title);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10418 2023-11-14 cs.LG cs.AI 81%

Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms

Seungju Han, Junhyeok Kim, Jack Hessel, Liwei Jiang, Jiwan Chung, Yejin Son, Yejin Choi, Youngjae Yu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at EMNLP 2023 (long)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04348 2023-11-09 cs.CL cs.AI 81%

Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning

Sai Munikoti, Anurag Acharya, Sridevi Wagle, Sameera Horawalavithana

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02216 2023-11-07 cs.CL cs.LG 81%

Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular Data

Mubashara Akhtar, Abhilash Shankarampeta, Vivek Gupta, Arpit Patil, Oana Cocarascu, Elena Simperl

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP 2023 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19301 2023-10-31 cs.CL cs.AI cs.CV 81%

ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense

Kankan Zhou, Eason Lai, Wei Bin Au Yeong, Kyriakos Mouratidis, Jing Jiang

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments This is the camera-ready version of the paper that will be published in the EMNLP 2023 Findings (Singapore, 6-10 December 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16755 2023-10-26 cs.CL cs.AI 81%

HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Yinghui He, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, Naihao Deng

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at Findings of EMNLP 2023

Journal ref Findings of EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14657 2023-10-24 cs.CL cs.AI 81%

Reasoning about Ambiguous Definite Descriptions

Stefan F. Schouten, Peter Bloem, Ilia Markov, Piek Vossen

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10632 2023-10-17 cs.CL cs.AI cs.RO 81%

BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology

Odhran O'Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud, Ali Essa Ghareeb, Justin Booth, Samuel G Rodriques

专题命中 推理评测 :planning(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023. Dataset and code: https://github.com/bioplanner/bioplanner

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02615 2023-10-16 cs.CL cs.AI 81%

How to Enhance Causal Discrimination of Utterances: A Case on Affective Reasoning

Hang Chen, Jing Luo, Xinyu Yang, Wenjing Zhu

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments accepted via EMNLP2023-main

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07018 2023-10-12 cs.CL cs.AI cs.RO 81%

NEWTON: Are Large Language Models Capable of Physical Reasoning?

Yi Ru Wang, Jiafei Duan, Dieter Fox, Siddhartha Srinivasa

专题命中 推理评测 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2023 Findings; 8 pages, 3 figures, 7 tables; Project page: https://newtonreasoning.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏