Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
基于边际的置信度排名用于可靠的LLM判断
Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang
机构
*
Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)
;
Institute of AI for Industries, Chinese Academy of Sciences(中国科学院工业人工智能研究所)
;
Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
ReplicatorBench:用于社会科学和行为科学中可复制性评估的LLM代理基准测试
Bang Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage, Zack Ranjan, Sai Koneru, Timothy M. Errington, Shakhlo Nematova, Sarah Rajtmajer, Jian Wu, Meng Jiang
机构
*
University of Notre Dame(圣母大学)
;
Old Dominion University(老道明大学)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
Independent Researcher(独立研究员)
;
Uppsala University(乌普萨拉大学)
;
Center for Open Science(开放科学中心)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
基于自我报告的LLM代理能够实现通用个体模拟
Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein
机构
*
Computer Science Department, Stanford University(斯坦福大学计算机科学系)
;
Department of Communication Studies, Northwestern University(西北大学传播学系)
;
Department of Communication, University of Washington(华盛顿大学传播学系)
;
Google DeepMind(谷歌DeepMind)
;
Department of Sociology, Stanford University(斯坦福大学社会学系)
;
Sciences Po(巴黎政治学院)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
LLM-ReSum: 一个通过自我评估实现LLM反思性总结的框架
Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Haihua Chen, Junhua Ding
机构
*
dept. of Information Science University of North Texas Denton, Texas, USA(信息科学系 俄克拉荷马州立大学 丹顿 俄克拉荷马州 美国)
;
dept. of Data Science University of North Texas Denton, Texas, USA(数据科学系 俄克拉荷马州立大学 丹顿 俄克拉荷马州 美国)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CommentsThis paper has been accepted as an invited paper for publication in Proceedings of The 12th IEEE International Conference on Big Data Computing Service and Machine Learning Applications. This is the accepted manuscript. The final authenticated version will be available via IEEE Xplore
机构
*
Department of Mathematics and Statistics(数学与统计学系)
;
Georgia State University(佐治亚州立大学)
;
Department of Probability and Statistics(概率与统计学系)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
Michigan State University(密歇根州立大学)
;
Concordia University(Concordia 大学)
;
Department of Statistics(统计学系)
;
University of Manitoba(曼尼托巴大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
GrowthHacker: 使用代码修改型LLM代理的自动离线策略评估优化
Jie JW Wu, Ayanda Patrick Herlihy, Ahmad Saleem Mirza, Ali Afoud, Fatemeh Fard
机构
*
Michigan Technological University, Houghton(密歇根技术大学)
;
Birmingham City University(伯明翰城市大学)
;
University of British Columbia, Kelowna(不列颠哥伦比亚大学, 肯洛纳)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
Evaluating LLM Personalization via Semantic Constraint Verification
通过语义约束验证评估LLM个性化
Xuran Li, Guanqin Zhang, Imran Razzak, Hakim Hacid, Eleanna Kafeza, Hao Xue, Flora D. Salim
机构
*
University of New South Wales(新南威尔士大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
The Technology Innovation Institute(技术创新研究所)
;
The Hong Kong University of Science and Technology(香港科技大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
衡量LLM导师是教学还是解题:教育影响的诊断方法
Junyi Yao, Zihao Zheng, Baichuan Li
机构
*
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Department of Operations Research and Engineering Management, Southern Methodist University(南卫理公会大学运筹学与工程管理系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation
VHDLSuite:面向LLM VHDL生成的统一流水线,包含数据合成与评估
Yijun Shen, Minghao Shao, Yichen Zhao, Zhuoyan Yu, Boyuan Chen, Yik-Cheung Tam, Muhammad Shafique
机构
*
Center for Data Science, NYU Shanghai, China(纽约市立大学上海分校数据科学中心)
;
NYU Tandon School of Engineering, USA(纽约大学Tandon工程学院)
;
NYU Abu Dhabi, UAE(纽约大学阿布扎比分校)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
评估LLM生成数据的质量与可信度综述
Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le, Fariha Kabir Torsha, Zhimeng Jiang, Minh Khai Bui, Chia-Yuan Chang, Yu-Neng Chuang, Zhen Xiong, Ying Lin, Guanchu Wang, Na Zou
机构
*
University of Houston(德克萨斯大学休斯敦分校)
;
Worcester Polytechnic Institute(沃思利理工学院)
;
Rice University(里德大学)
;
Texas A&M University(德克萨斯农工大学)
;
University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
;
University of Southern California(南加州大学)
;
University of North Carolina at Charlotte(北卡罗来纳州立大学夏洛特分校)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
分析LLM生成文本中说服性语言的差异:揭示刻板的性别模式
Amalie Brogaard Pauli, Maria Barrett, Max Müller-Eberstein, Isabelle Augenstein, Ira Assent
机构
*
Department of Computer Science, Aarhus University(阿arhus大学计算机科学系)
;
AMD Silo AI
;
University of Tokyo(东京大学)
;
IT University of Copenhagen(哥本哈根IT大学)
;
Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment
从评分到解释:评估基于量规的教学质量评估中的SHAP和LLM理由
Ivo Bueno, Babette Bühler, Philipp Stark, Tim Fütterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Lund University(吕勒奥大学)
;
University of Tübingen(图宾根大学)
;
Stanford Graduate School of Education(斯坦福大学教育研究生院)
;
Harvard Graduate School of Education(哈佛大学教育研究生院)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
一致且独特:基于相似图最大独立集提示选择的LLM基准测试效率
Denica Kjorvezir, Marko Djukanović, Ana Gjorgjevikj, Gjorgjina Cenikj, Tome Eftimov
机构
*
Computer Systems Department, Jožef Stefan Institute, Ljubljana, Slovenia(计算机系统部,乔塞夫·斯塔芬研究所,卢布尔雅那,斯洛文尼亚)
;
Jožef Stefan International Postgraduate School, Ljubljana, Slovenia(乔塞夫·斯塔芬国际研究生学院,卢布尔雅那,斯洛文尼亚)
;
Center for Astrophysics and Cosmology, University of Nova Gorica, Nova Gorica, Slovenia(天体物理与宇宙学中心,诺瓦戈里察大学,诺瓦戈里察,斯洛文尼亚)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI