机构
*
National University of Singapore(新加坡国立大学)
;
Yunnan University(云南大学)
;
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
Laboratory for Statistical Monitoring and Intelligent Governance of Common Prosperity, Zhejiang Gongshang University(浙江工商大学共同富裕统计监测与智能治理实验室)
;
Tongyi Lab, Alibaba Group(阿里集团通义实验室)
;
Alibaba Cloud(阿里云)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
通过项目反应理论诊断LLM作为评判者的可靠性
Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, Bugeun Kim
机构
*
Department of Artificial Intelligence, Chung-Ang University, Seoul, Republic of Korea(Chung-Ang 大学人工智能系)
;
Department of Industrial Engineering, Seoul National University, Seoul, Republic of Korea(首尔国立大学工业工程系)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
基于LLM的自动评分中可学习的评估技能:通过迭代优化构建评分标准
Yun Wang, Xin Xia, Xuansheng Wu, Xiaoming Zhai, Ninghao Liu
机构
*
School of Computing, University of Georgia, Athens, GA, USA(佐治亚大学计算机学院)
;
AI4STEM Education Center, University of Georgia, Athens, GA, USA(佐治亚大学AI4STEM教育中心)
;
The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学)
机构
*
AI Lab, Princeton University Engineering and Applied Sciences(普林斯顿大学人工智能实验室、工程与应用科学学院)
;
RAND Corporation(RAND公司)
;
Engineering and Applied Sciences(工程与应用科学)
;
Science, Technology, and International Affairs(科学、技术与国际事务)
;
Georgetown University(乔治·华盛顿大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints
从学习资源到能力:基于证据和图约束的LLM标签方法
Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge
机构
*
Université de technologie de Compiègne, CNRS, Heudiasyc(法国图卢兹技术大学、CNRS、Heudiasyc实验室)
;
Sorbonne Université, CNRS UMR 7585, LPMHE(巴黎大学、CNRS UMR 7585、LPMHE实验室)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents
DynaSchedBench: 基于LLM的调度代理中的校准动态调度基准与可观测性悖论
Shijie Cao, Yuan Yuan, Jing Liu
机构
*
School of Computer Science and Engineering, Beihang University, Beijing 100191, China(北航计算机科学与工程学院)
;
Shenzhen Loop Area Institute, Shenzhen, China(深圳环城院)
;
Qingdao Research Institute, Beihang University(北航青岛研究院)
;
Hangzhou Innovation Institute, Beihang University(北航杭州创新院)
;
School of Artificial Intelligence, Xidian University, Xi'an 710071, Shaanxi, China(西电人工智能学院)
;
Guangzhou Institute of Technology, Xidian University, Guangzhou 510555, Guangdong, China(西电广州技术院)
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
University of Southern California(南加州大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
;
University of Arizona(亚利桑那大学)