机构
*
City University of Hong Kong(香港城市大学)
;
Southern University of Science and Technology(南方科技大学)
;
Xi’an Jiaotong University(西安交通大学)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
Cognitive Models and AI Algorithms Provide Templates for Designing Language Agents
认知模型和AI算法为设计语言代理提供模板
Ryan Liu, Dilip Arumugam, Cedegao E. Zhang, Sean Escola, Xaq Pitkow, Thomas L. Griffiths
机构
*
Department of Computer Science, Princeton University(普林斯顿大学计算机科学系)
;
Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology(麻省理工学院脑科学与认知科学系)
;
Zuckerman Mind Brain Behavior Institute(祖克曼心智大脑行为研究所)
;
Department of Psychiatry, Columbia University(哥伦比亚大学精神医学系)
;
Neuroscience Institute(神经科学研究所)
;
Department of Machine Learning, Carnegie Mellon University(卡内基梅隆大学机器学习系)
;
Department of Psychology, Princeton University(普林斯顿大学心理学系)
专题命中
评测与基准
:language agent(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)
机构
*
Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空信息安全与可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学)
;
Department of Computer Science and Engineering, University at Buffalo, SUNY(计算机科学与工程系、布法罗大学、纽约州立大学)
;
Department of Computer Science, College of Information Science and Technology, Jinan University(计算机科学系、信息科学与技术学院、济南大学)
;
College of Cyber Security, Jinan University(网络安全学院、济南大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation
SC-Arena: 一种用于具有知识增强评估的单细胞推理的自然语言基准
Jiahao Zhao, Feng Jiang, Shaowei Qin, Zhonghui Zhang, Junhao Liu, Guibing Guo, Hamid Alinejad-Rokny, Min Yang
机构
*
Software College, Northeastern University(东北大学软件学院)
;
Shenzhen University of Advanced Technology(深圳大学先进技术学院)
;
Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院)
;
University of California, Irvine(加州大学 Irvine 分校)
;
School of Biomedical Engineering, UNSW Sydney(悉尼大学生物医学工程学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI
Journal refPublished at 11th Language & Technology Conference: Human Language Technologies as a Challenge for Computer Science, Linguistics and Low Resourced Languages, 2025, Poznań, Poland
CommentsSubmitted to CogSci 2025; see more at https://jmuchovej.com/projects/llm-tom. Note: "abstractness" is the second feature we test for, but due to arXiv's abstract requirements, the text has been altered
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
PoSh:利用场景图引导LLM-as-a-Judge进行详细图像描述
Amith Ananthram, Elias Stengel-Eskin, Lorena A. Bradford, Julia Demarest, Adam Purvis, Keith Krut, Robert Stein, Rina Elster Pantalony, Mohit Bansal, Kathleen McKeown
机构
*
Columbia University(哥伦比亚大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
The National Gallery of Art(国家艺术馆)
;
UCLA(加州大学洛杉矶分校)
;
UNC Chapel Hill(北卡罗来纳大学教堂山分校)
机构
*
Yau Mathematical Sciences Center, Tsinghua University(清华大学尤洋数学科学中心)
;
Iluvatar CoreX
;
Department of Mathematics, Southern University of Science and Technology(南方科技大学数学系)
;
Westlake Institute for Advanced Study, Westlake University(西湖研究学院)
;
Qiuzhen College, Tsinghua University(清华大学齐臻学院)
;
Department of Physics, The Chinese University of Hong Kong(香港中文大学物理系)
;
Yanqi Lake Beijing Institute of Mathematical Sciences and Applications (BIMSA)(燕琦湖北京应用数学研究所(BIMSA))
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG