Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War
LLM时代:战争迷雾下大型语言模型的推理、外交与可靠性的战略1v1基准测试
Arnaud Ricci
机构
*
Independent researcher, Switzerland(瑞士独立研究员)
专题命中
推理与问题求解
:LLM(title,title_cn);large language model(title);language model(title);分类 cs.CL、cs.AI
AI总结
提出Age of LLM基准,在13x7网格上让两个LLM进行1v1对战,包含战争迷雾、完全外交和严格JSON格式可靠性维度,通过54场比赛评估15个推理模型,发现核冲刺主导、外交频繁但极少达成、非法动作反映信念追踪,并初步关联可靠性与获胜。
Comments25 pages including appendices, 8 figures, 4 tables; appendices include verbatim system prompt and engine resolution pseudocode. All correlations reported with p-values, 95% bootstrap confidence intervals and Spearman's rho; includes a Steiger test and Bradley-Terry fit
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
SICI:一种揭示LLM立场检测中相变的语义-语用复杂度指数
Fuqiang Niu, Bowen Zhang
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络空间安全学院)
;
School of Artificial Intelligence, Shenzhen Technology University(深圳技术大学人工智能学院)
A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial
一种用于加速罕见病诊断的专用推理大语言模型:随机AI辅助医生试验
Haichao Chen, Songchi Zhou, Zhengyun Zhao, Shikai Hu, Xianghong Jin, Hongwei Ji, Li He, Shuli Li, Yiming Qin, Xin Tan, Runfeng Shi, Yih Chung Tham, Jiaye Zhu, Ye Li, Ye Jin, Longhao Cao, Dawei Li, Honghan Wu, Hongqiu Gu, Guanqiao Li, Tudor Groza, Chunying Li, Dian Zeng, Weihong Yu, Gareth Baynam, Saumya Shekhar Jamuar, Min Shen, Shuyang Zhang, Bin Sheng, Sheng Yu, Tien Yin Wong
机构
*
Tsinghua Medicine, Tsinghua University(清华大学医学部,清华大学)
;
Department of Ophthalmology, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College(北京大学人民医院眼科,中国医学科学院 & 北京大学医学部)
;
Department of Statistics and Data Science, Tsinghua University(清华大学统计与数据科学系)
;
Department of Rheumatology and Clinical Immunology, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College(北京大学人民医院风湿免疫科,中国医学科学院 & 北京大学医学部)
;
Department of Rare Diseases, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College(北京大学人民医院罕见病科,中国医学科学院 & 北京大学医学部)
;
Department of Dermatology, Xijing Hospital, Air Force Medical University(西安空军军医大学西京医院皮肤科)
;
School of Computer Science and Technology, East China Normal University (ECNU)(东华大学计算机科学与技术学院)
;
Department of Neurosurgery, Xuanwu Hospital, Capital Medical University(首都医科大学宣武医院神经外科)
;
Department of Ophthalmology, National University of Singapore(新加坡国立大学眼科)
;
Singapore Eye Research Institute, Singapore National Eye Centre(新加坡眼科学研究所,新加坡国家眼科中心)
;
Ophthalmology and Visual Science Academic Clinical Program, Duke-NUS Medical School(杜克-国立新加坡大学医学学校眼科与视觉科学学术临床项目)
;
Department of Pediatrics, The First Hospital of Tsinghua University(清华大学第一医院儿科)
;
College of Future Technology, Peking University(北京大学未来技术学院)
;
School of Health and Wellbeing, University of Glasgow(格拉斯哥大学健康与福祉学院)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Beihang University(北京航空航天大学)
;
Simon Fraser University(西蒙弗雷泽大学)
专题命中
推理与问题求解
:LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
VeryTrace: 通过可编译形式化与结构化验证验证推理轨迹
Ninghan Zhong, Ahmet Ege Tanriverdi, Kaan Kale, Sriram Vishwanath
机构
*
School of Electrical and Computer Engineering, Georgia Institute of Technology, USA(佐治亚理工学院电子与计算机工程学院)
;
Department of Electrical and Computer Engineering, Bogazici University, Turkey(博亚奇大学电子与计算机工程系)
机构
*
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学基础模型研究所)
;
School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院)
专题命中
推理与问题求解
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
HANCLIP:双曲角否定视觉语言模型系列
Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen, Liting Zhou, Cathal Gurrin
机构
*
ADAPT Centre Dublin City University, Ireland(爱尔兰都柏林城市大学ADAPT中心)
;
University of East Anglia Norwich, UK(英国东英吉利大学)
;
Queen’s University Belfast Belfast, UK(英国贝尔法斯特女王大学)
;
University of Science Vietnam National University Ho Chi Minh City, Vietnam(越南胡志明市国家大学理科大学)
机构
*
State Key Laboratory of Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
;
Beijing Academy of Artificial Intelligence(北京智源人工智能研究院)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
YiXin-AILab, YIXIN(YiXin-AILab,YIXIN)
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
AfriqueLLM: 数据混合与模型架构如何影响非洲语言的持续预训练
Hao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani
机构
*
McGill University(麦吉尔大学)
;
Mila-Quebec AI Institute(魁北克AI研究所)
;
LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)
;
Microsoft AI for Good Research Lab(微软AI for Good研究实验室)
专题命中
推理与问题求解
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL