Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
基于理论的LLM推理中人类似 fallacy 模式的评估
Andrew Keenan Richardson, Ryan Othniel Kearns, Sean Moss, Vincent Wang-Mascianica, Philipp Koralus
机构
*
The Laboratory for Human-Centered AI, Institute for Ethics in AI, Faculty of Philosophy, University of Oxford, UK(以人为中心的人工智能实验室,人工智能伦理研究所,哲学学院,牛津大学,英国)
;
Oxford Internet Institute, University of Oxford, UK(牛津互联网研究所,牛津大学,英国)
;
School of Computer Science, University of Birmingham, UK(计算机科学学院,伯明翰大学,英国)
SemEval-2026 Task 12: Abductive Event Reasoning: Towards Real-World Event Causal Inference for Large Language Models
SemEval-2026任务12:归纳事件推理:面向大规模语言模型的现实事件因果推断
Pengfei Cao, Mingxuan Yang, Yubo Chen, Chenlong Zhang, Mingxuan Liu, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国)
机构
*
East China Normal University(东华大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shanghai Innovation Institute(上海创新研究院)
;
University of Agder(阿格德大学)
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
GIR-Bench:用于生成图像的多功能基准
Hongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu, Yuhang Yang, Xiaoshuang Huang, Jiayin Cai, Xiaolong Jiang, Yao Hu, Long Chen
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Peking University(北京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Xiaohongshu Inc.(小红书公司)
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
Patho-R1: 基于多模态强化学习的病理专家推理器
Wenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu
机构
*
Department of Pathology, West China Hospital, Sichuan University(四川大学华西医院病理科部门)
;
Institute of Clinical Pathology, West China Hospital, Sichuan University(四川大学华西医院临床病理科研究所)
;
University of Toronto(多伦多大学)
;
Business School, Sichuan University(四川大学商学院)
;
Department of Pathology, Shengjing Hospital of China Medical University(中国医科大学盛京医院病理科部门)
Comments39 pages, 5 figures, 5 tables. Preprint. Submitted to NIST CAISI (Docket NIST-2025-0035, March 2026). Also available on Zenodo: https://doi.org/10.5281/zenodo.18971110
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
阿拉伯方言MMLU:评估阿拉伯语和多语言模型在阿拉伯方言上的能力
Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih, Kirill Chirkunov, Muhammed AbuOdeh, Radu Florian, Teresa Lynn, Preslav Nakov, Alham Fikri Aji
机构
*
IBM Research AI(IBM人工智能研究院)
;
New York University Abu Dhabi(纽约大学迪拜分校)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
SVBRD-LLM: Self-Verifying Behavioral Rule Discovery for Autonomous Vehicle Identification
SVBRD-LLM:自主车辆识别的自验证行为规则发现
Xiangyu Li, Tianyi Wang, Junfeng Jiao, Christian Claudel, Zhaomiao Guo
机构
*
Fariborz Maseeh Department of Civil, Architectural, and Environmental Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校土木、建筑与环境工程学院)
;
School of Architecture, The University of Texas at Austin(德克萨斯大学奥斯汀分校建筑学院)
机构
*
Independent Researcher(独立研究者)
;
National Institute for Data Science in Health and Medicine, Xiamen University, Xiamen, China(健康医学数据科学国家研究院,厦门大学,厦门,中国)
;
Yue’erwan Internet Hospital Co., Ltd.(悦尔湾互联网医院有限公司)
;
School of Biomedical Engineering, Tsinghua University, Beijing, China(生物医学工程学院,清华大学,北京,中国)
;
National University of Singapore, Singapore(新加坡国立大学)
;
Yttrium-90 Precision Interventional Radiotherapy Center of Liver Cancer, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua University, Beijing, China(肝癌钇-90精准介入放射治疗中心,北京清华大学昌平医院,临床医学学院,清华大学,北京,中国)
;
Hepatobiliary and Pancreatic Centre, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua University, Beijing, China(肝胆胰中心,北京清华大学昌平医院,临床医学学院,清华大学,北京,中国)