CommentsAccepted at the ICML 2026 workshops "Statistical Frameworks for Uncertainty in Agentic Systems" and "Combining Theory and Benchmarks: Towards a Virtuous Cycle to Understand and Guarantee Foundation Model Performance". 13 pages, 9 figures
Culturally uneven urban perception in large language models
大型语言模型通过文化不平等的基线感知城市
Rong Zhao, Wanqi Liu, Zhizhou Sha, Nanxi Su, Yecheng Zhang, Ying Long
机构
*
Centre for Advanced Spatial Analysis (CASA), UCL, London, UK(高级空间分析中心(CASA),伦敦大学学院,英国)
;
School of Architecture, Tsinghua University, Beijing, China(清华大学建筑学院,北京,中国)
;
Department of Computer Science, UT Austin, Austin, TX, USA(得克萨斯大学奥斯汀分校计算机科学系,奥斯汀,德克萨斯,美国)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);prompting(abstract)
CommentsICAHS, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
PeruMedQA:在秘鲁医学考试上评估大语言模型(LLMs)——数据集构建与评估
Rodrigo M. Carrillo-Larco, Jesus Lovón Melgarejo, Manuel Castillo-Cara, Gusseppe Bravo-Rocca
机构
*
Hubert Department of Global Health, Rollins School of Public Health, Emory University(霍伯特全球健康部门,埃默里大学公共卫生学院)
;
Emory Global Diabetes Research Center of Woodruff Health Sciences Center, Emory University(埃默里大学伍德鲁夫健康科学中心全球糖尿病研究中心)
;
Institut de Recherche en Informatique de Toulouse(图卢兹信息研究院)
;
Universidad Nacional de Educación a Distancia(远程教育国立大学)
;
Instituto de Investigación Científica, Universidad de Lima(科学研究所,利马大学)
;
Barcelona Supercomputing Center(巴塞罗那超级计算中心)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG
Integrating Virtual Reality and Large Language Models for Team-Based Non-Technical Skills Training and Evaluation in the Operating Room
将虚拟现实与大型语言模型结合用于手术室基于团队的非技术技能训练与评估
Jacob Barker, Doga Demirel, Cullen Jackson, Anna Johansson, Robbin Miraglia, Darian Hoagland, Stephanie B. Jones, John Mitchell, Daniel B. Jones, Suvranu De
机构
*
Beth Israel Deaconess Medical Center Center(贝希斯尔德医疗中心中心)
;
Department of Surgery, Northwell Health(外科,北well健康)
;
College of Engineering, Florida Agricultural and Mechanical University and Florida State University(工程学院,佛罗里达农业与机械大学和佛罗里达州立大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
机构
*
Baidu Inc.(百度公司)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhejiang University(浙江大学)
;
University of Technology Sydney(悉尼科技大学)
GhazalBench: Evaluating LLM Understanding and Canonical Surface-Form Access in Persian Ghazals
GhazalBench: 评估大语言模型对波斯抒情诗的理解与规范表层形式访问
Ghazal Kalhor, Yadollah Yaghoobzadeh
机构
*
School of Electrical and Computer Engineering, College of Engineering, University of Tehran(德黑兰理工大学电气与计算机工程学院)
;
Tehran Institute for Advanced Studies, Khatam University(德黑兰高级研究院,凯塔姆大学)
专题命中
评测与基准
:LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL
机构
*
AnnLab(安实验室)
;
Institute of Semiconductors, Chinese Academy of Sciences(中国科学院半导体研究所)
;
Zhongguancun Academy(中关村学院)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
State Key Laboratory of High Performance Ceramics(高性能陶瓷国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Electronic, Electrical and Communication Engineering(电子电气与通信工程学院)
;
University of ChineseAcademy of Sciences(中国科学院大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
学习承诺:通过学习异步解码扩展语言模型解码并行性
Tian Jin, Ellie Y. Cheng, Zack Ankner, Nikunj Saunshi, Blake M. Elias, Amir Yazdanbakhsh, Jonathan Ragan-Kelley, Suvinay Subramanian, Michael Carbin
机构
*
DeepMind, London, UK(深度思维公司,伦敦,英国)
;
Google Research, New York, NY, USA(谷歌研究院,纽约,纽约州,美国)
;
Stanford University, Stanford, CA, USA(斯坦福大学,斯坦福,加利福尼亚州,美国)
;
University of Toronto, Toronto, Ontario, Canada(多伦多大学,多伦多,安大略省,加拿大)
;
University of Washington, Seattle, WA, USA(华盛顿大学,西雅图,华盛顿州,美国)
专题命中
评测与基准
:language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);分类 cs.CL、cs.LG
机构
*
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
East China Normal University(华东师范大学)
;
New York University(纽约大学)
;
Tongji University(同济大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS
FEniCS中变分多物理场有限元模拟的受约束自然语言接口
Nilay Upadhyay, Wesley F. Reinhart
机构
*
Department of Engineering Science and Mechanics, The Pennsylvania State University(工程科学与力学系,宾夕法尼亚州立大学)
;
Department of Materials Science and Engineering, The Pennsylvania State University(材料科学与工程系,宾夕法尼亚州立大学)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG