DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
阿拉伯方言MMLU:评估阿拉伯语和多语言模型在阿拉伯方言上的能力
Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih, Kirill Chirkunov, Muhammed AbuOdeh, Radu Florian, Teresa Lynn, Preslav Nakov, Alham Fikri Aji
机构
*
IBM Research AI(IBM人工智能研究院)
;
New York University Abu Dhabi(纽约大学迪拜分校)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
URAG: A Benchmark for Uncertainty Quantification in Retrieval-Augmented Large Language Models
URAG:一种用于检索增强大型语言模型不确定性量化的基准
Vinh Nguyen, Cuong Dang, Jiahao Zhang, Hoa Tran, Minh Tran, Trinh Chau, Thai Le, Lu Cheng, Suhang Wang
机构
*
Uppsala University(乌普萨拉大学)
;
University of Science, VNU-HCM(VNU-HCM大学)
;
Indiana University(印第安纳大学)
;
FPT Software, AI Center(FPT软件人工智能中心)
;
The Pennsylvania State University(宾夕法尼亚州立大学)
;
VNU University of Engineering(VNU工程大学)
;
University of Illinois at Chicago(伊利诺伊大学芝加哥分校)
Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm
大型语言模型具备心智理论吗?使用奇怪故事范式进行比较评估
Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
机构
*
ELTE Research Centre for Linguistics(ELTE语言学研究中心)
;
Department of Cognitive Science, Faculty of Natural Sciences, Budapest University of Technology and Economics(布达佩斯技术与经济大学自然科学学院认知科学系)
;
Institute of Mathematics, Eötvös Lóránd University(欧多布雷尼大学数学研究所)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
SWE-QA-Pro:一个代表性的基准和可扩展的训练配方用于仓库级代码理解
Songcheng Cai, Zhiheng Lyu, Yuansheng Ni, Xiangchao Chen, Baichuan Zhou, Shenzhe Zhu, Yi Lu, Haozhe Wang, Chi Ruan, Benjamin Schneider, Weixu Zhang, Xiang Li, Andy Zheng, Yuyu Zhang, Ping Nie, Wenhu Chen
机构
*
University of Waterloo(滑铁卢大学)
;
University of Toronto(多伦多大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
McGill University & MILA(麦吉尔大学及MILA)
;
Verdent AI, Inc.(Verdent AI公司)
AOI: Turning Failed Trajectories into Training Signals for Autonomous Cloud Diagnosis
AOI: 将失败轨迹转化为自主云诊断的训练信号
Pei Yang, Wanyi Chen, Asuka Yuxi Zheng, Xueqian Li, Xiang Li, Haoqin Tu, Jie Xiao, Yifan Pang, Dongdong Zhang, Fuqiang Li, Alfred Long, Lynn Ai, Eric Yang, Bill Shi
机构
*
Gradient
;
Soochow University(苏州大学)
;
UC Santa Cruz
;
Georgia Institute of Technology(佐治亚理工学院)
;
University College London(伦敦大学学院)
;
WeJoy
;
ByteDance(字节跳动)
ReGAIN: Retrieval-Grounded AI Framework for Network Traffic Analysis
ReGAIN:基于检索的网络流量分析人工智能框架
Shaghayegh Shajarian, Kennedy Marsh, James Benson, Sajad Khorsandroo, Mahmoud Abdelsalam
机构
*
Computer Science(计算机科学)
;
North Carolina A\&T State University(北卡罗来纳A&T州立大学)
;
Institute for Cyber Security(网络安全研究所)
;
University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)
CommentsAccepted for publication in 20th International Conference on Agents and Multi-Agent Systems: Technologies and Applications (AMSTA 2026), to appear in Springer Nature proceedings (KES Smart Innovation Systems and Technologies). The final authenticated version will be available online at Springer
AutoClimDS: Climate Data Science Agentic AI -- A Knowledge Graph is All You Need
AutoClimDS:气候数据科学代理AI——一个知识图谱足矣
Ahmed Jaber, Wangshu Zhu, Ayon Roy, Karthick Jayavelu, Justin Downes, Sameer Mohamed, Candace Agonafir, Linnia Hawkins, Tian Zheng
机构
*
NSF STC Learning the Earth with AI and Physics (LEAP), Columbia University(NSF STC 学习地球与人工智能和物理(LEAP),哥伦比亚大学)
;
AWS Generative AI Innovation Center(AWS 生成式人工智能创新中心)
;
Department of Statistics, Columbia University(哥伦比亚大学统计系)
机构
*
Department of Pathology, University of Yamanashi, Chuo, Japan(山梨大学病理科)
;
Division of Pathology, Exploratory Oncology Research & Clinical Trial Center, National Cancer Center, Kashiwa, Japan(国立癌症中心探索肿瘤研究与临床试验中心病理科)
;
Department of Preventive Medicine, Graduate School of Medicine, The University of Tokyo, Tokyo, Japan(东京大学医学部预防医学科)
;
Department of Medical Oncology, National Cancer Center Hospital East, Kashiwa, Japan(国立癌症中心东医院医学肿瘤科)
;
Department of Thoracic Surgery, National Cancer Center Hospital East, Kashiwa, Japan(国立癌症中心东医院胸外科)