机构
*
Tsinghua University(清华大学)
;
SenseTime Research(商汤科技研究院)
;
Peking University(北京大学)
;
Technical University of Munich, Heilbronn Data Science Center, Munich Data Science Institute(慕尼黑技术大学,海德堡数据科学中心,慕尼黑数据科学研究所)
机构
*
Department of Pathology, University of Yamanashi, Chuo, Japan(山梨大学病理科)
;
Division of Pathology, Exploratory Oncology Research & Clinical Trial Center, National Cancer Center, Kashiwa, Japan(国立癌症中心探索肿瘤研究与临床试验中心病理科)
;
Department of Preventive Medicine, Graduate School of Medicine, The University of Tokyo, Tokyo, Japan(东京大学医学部预防医学科)
;
Department of Medical Oncology, National Cancer Center Hospital East, Kashiwa, Japan(国立癌症中心东医院医学肿瘤科)
;
Department of Thoracic Surgery, National Cancer Center Hospital East, Kashiwa, Japan(国立癌症中心东医院胸外科)
CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?
CoMMET:大型语言模型在理论思维任务中能发挥多大作用?
Ruirui Chen, Weifeng Jiang, Chengwei Qin, Cheston Tan
机构
*
Agency for Science, Technology and Research (A*STAR)(科技研究局)
;
Nanyang Technological University(南洋理工大学)
;
Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州),中国)
MaterialFigBENCH: benchmark dataset with figures for evaluating college-level materials science problem-solving abilities of multimodal large language models
MaterialFigBENCH:用于评估多模态大语言模型在大学级材料科学问题解决能力的基准数据集
Michiko Yoshitake, Yuta Suzuki, Ryo Igarashi, Yoshitaka Ushiku, Keisuke Nagato