Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification
验证你的权威:在多标签先例处理分类上对LLM进行基准测试
M. Mikail Demir, M. Abdullah Canbaz
机构
*
Department of Information Science and Technology(信息科学与技术系)
;
College of Emergency Preparedness, Homeland Security, and Cybersecurity(应急准备、国土安全与网络安全学院)
;
University at Albany, SUNY(萨利纳大学)
QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI
QQJ: 量化定性判断以实现可扩展且与人类对齐的生成AI评估
Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare, Mohammad Sabouri
机构
*
AI Lab, Arioobarzan Engineering Team(艾伊罗巴赞工程团队人工智能实验室)
;
Department of Computer Engineering and Information Technology(计算机工程与信息科技系)
;
Department of Informatics, Bioengineering, Robotics and Systems Engineering(信息学、生物工程、机器人与系统工程系)
;
University of Genoa(热那亚大学)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
ADR:一种用于企业代理AI安全的代理检测系统
Chenning Li, Pan Hu, Justin Xu, Baris Ozbas, Olivia Liu, Caroline Van, Manxue Li, Wei Zhou, Mohammad Alizadeh, Pengyu Zhang, KK Sriramadhesikan, Ming Zhang
机构
*
Department of Automation, Tsinghua University, Beijing, China(清华大学自动化系)
;
Alibaba International Digital Commerce Group, Beijing, China(阿里巴巴国际数字 commerce 集团)
;
School of Software, Tsinghua University, Beijing, China(清华大学软件学院)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
PluRule:一种用于社交媒体上多元社区调节的基准测试
Zoher Kachwala, Bao Tran Truong, Rasika Muralidharan, Haewoon Kwak, Jisun An, Filippo Menczer
机构
*
Observatory on Social Media, Indiana University, USA(社交媒体观察站,印第安纳大学,美国)
;
Center Synergy of Systems, TUD Dresden University of Technology, Germany(系统协同中心,德累斯顿技术大学,德国)
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
CoCoReviewBench:面向AI审稿人完整性和正确性的基准测试
Hexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu, Dehao Huang, Derek F. Wong, Yue Wang, Xuebo Liu, Min Zhang
机构
*
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳研究院)
;
Zhongguancun Academy, Beijing, China(中关村学院)
;
Faculty of Computing, Harbin Institute of Technology, Harbin, China(哈尔滨工业大学计算机学院)
;
Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen, China(南方科技大学计算机科学与工程系)
;
NLP²CT Lab, Department of Computer and Information Science, University of Macau, China(澳门大学自然语言处理实验室)
Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes
帮助陷入困境的客户:一个基于LLM的代理,能够对话、探测和分流
Alankar Atreya, Stefan Sylvius Wanger, Devesh Batra, Robert Hankache, Cristovao Iglesias, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
机构
*
South China University of Technology(华南理工大学)
;
Tencent Financial Technology(腾讯金融科技)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Soochow University(苏州大学)
;
Zhejiang Key Laboratory of Intelligent Education Technology and Application, Zhejiang Normal University(浙江省智能教育技术与应用重点实验室,浙江师范大学)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
RxEval: 一个处方级基准,用于评估LLM药物推荐
Shuhao Chen, Weisen Jiang, Changmiao Wang, Xiaoqing Wu, Xuanren Shi, Yu Zhang, James T. Kwok
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
Southern University of Science and Technology(南方科技大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区)
;
Shenzhen University General Hospital(深圳大学人民医院)
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction
Collider-Bench:通过粒子物理分析复现评估AI代理
Darius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang, Siddharth Mishra-Sharma, David Shih
机构
*
New High Energy Theory Center(新高能理论中心)
;
Department of Physics & Astronomy(物理与天文学系)
;
Rutgers University(罗格斯大学)
;
Faculty of Computing & Data Sciences(计算与数据科学学院)
Rethinking Efficient Graph Coarsening via a Non-Selfishness Principle
重新思考通过非自私原则的高效图粗化
Xu Bai, Bin Lu, Kun Zhang, Shengbo Chen, Xinbing Wang, Chenghu Zhou, Meng Jin
机构
*
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China(上海交通大学信息科学与电子工程学院)
;
School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院)
;
School of Environment Science and Engineering, Shanghai Jiao Tong University, Shanghai, China(上海交通大学环境科学与工程学院)
;
School of Artificial Intelligence, Nanchang University, Nanchang, China(南昌大学人工智能学院)
;
Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing, China(中国科学院地理科学与资源研究所)
Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
为高能物理中的中微子事件分类适应视觉-语言模型
Dikshant Sagar, Kaiwen Yu, Alejandro Yankelevich, Jianming Bian, Pierre Baldi
机构
*
Department of Computer Science, University of California, Irvine, CA, USA(计算机科学系,加州大学欧文分校,加州,美国)
;
Department of Physics, University of California, Irvine, CA, USA(物理系,加州大学欧文分校,加州,美国)
PERCEIVE: A Benchmark for Personalized Emotion and Communication Behavior Understanding on Social Media
PERCEIVE:面向社交媒体个性化情绪与交流行为理解的基准测试
Jian Liao, Yujin Zheng, Suge Wang, Jianxing Zheng, Deyu Li
机构
*
School of Computer and Information Technology, Shanxi University, China(山西大学计算机与信息学院)
;
Key Laboratory of Computational Intelligence and Chinese Information Processing of Ministry of Education, Shanxi University, China(教育部计算智能与中文信息处理重点实验室,山西大学,中国)
;
Joint Laboratory of Tourism Big Data in Shanxi Province, China(山西省旅游大数据联合实验室)