UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
UrbanWell: 面向时空城市福祉分析的多模态大语言模型基准测试
Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
机构
*
University of Helsinki(赫尔辛基大学)
;
Zhongguancun Academy(中关村学院)
;
University of Oxford(牛津大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
DLawBench: 通过多轮法律咨询评估大语言模型
Li Zhang, Yuzhen Shi, Yiran Hu, Jingwen Zhang, Wenbo Lv, Yubo Ma, Wei Wang, Rongyao Shi, Yuanyang Qiu, Xinran Xu, Yuemeng Qi, Linlin Miao, Jaromir Savelka, Yun Liu, Kevin Ashley, Bing Zhao, Hu Wei, Lin Qu
机构
*
University of Pittsburgh(匹兹堡大学)
;
Alibaba Group(阿里巴巴集团)
;
University of Waterloo(滑铁卢大学)
;
Shandong University(山东大学)
;
Skylenage
;
China University of Political Science and Law(中国政法大学)
;
Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)
;
Fordham University(福特汉姆大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Tsinghua University(清华大学)
Constrained Semantic Decompression in LLMs through Persian Proverb-Conditioned Story Generation
通过波斯谚语条件故事生成实现LLM中的约束语义解压缩
Zahra Habibzadeh, Paria Khoshtab, Amir Mesbah, Yadollah Yaghoobzadeh
机构
*
Tehran Institute for Advanced Studies, Khatam University, Iran(德黑兰高级研究所,卡塔姆大学,伊朗)
;
School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran(电气与计算机工程学院,工程学院,德黑兰大学,伊朗)
Comments65 pages, 3 figures, 5 tables. Reference architecture with a reference implementation of the policy-engine core and microbenchmark results; full-system evaluation identified as future work
Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
重新评估高性能大语言模型在波兰医学考试中的表现:真实能力还是偏差驱动?
Antoni Lasik, Jakub Pokrywka, Łukasz Grzybowski, Jeremi Ignacy Kaczmarek, Gabriela Korzańska, Janusz Świeczkowski-Feiz, Oskar Pastuszek, Paulina Hoffman, Jakub Tomasz Dąbrowski, Wojciech Kusa
机构
*
NASK National Research Institute(NASK国家研究所)
;
Adam Mickiewicz University(亚当·密茨凯维奇大学)
;
ARAAI Poland(ARAAI波兰)
;
Poznań University of Medical Sciences(波兹南医科大学)
;
Centre of Postgraduate Medical Education, Poland(波兰研究生医学教育中心)
;
T. Marciniak Lower Silesian Specialist Hospital(T. 马尔奇尼亚克下西里西亚专科医院)
;
Medical University of Warsaw(华沙医科大学)
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
RAIL: 基于CHC框架重新思考大型音频语言模型中的听觉智能
Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang
机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
Faculty of Psychology and Educational Sciences, Alexandru Ioan Cuza University of Iași(亚历山德鲁伊万库扎大学心理学与教育科学学院)
;
School of Electronic Information, Wuhan University(武汉大学电子信息学院)
;
School of Public Health, The University of Hong Kong(香港大学公共卫生学院)
;
School of Computer Science, The University of Auckland(奥克兰大学计算机科学学院)
;
Department of Data Science and Artificial Intelligence, Monash University(莫纳什大学数据科学与人工智能系)