REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
REAL: 面向LLM评判的回归感知强化学习
Yasi Zhang, Tianyu Chen, Mingyuan Zhou, Oscar Leong, Ying Nian Wu, Michal Lukasik
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
The University of Texas at Austin(得克萨斯大学奥斯汀分校)
;
Google Research Now at Google DeepMind(谷歌研究 现在在谷歌深Mind)
专题命中
评测与基准
:LLM(title,title_cn);SFT(summary_cn,abstract);large language model(abstract);language model(abstract)
机构
*
National Institute of Informatics, Japan(日本国立信息研究所)
;
The University of Tokyo, Japan(东京大学)
;
Inria, LS2N, Nantes Université, France(法国Inria、LS2N、南特大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
CommentsSubmitted to ACM Journal on Responsible Computing, Special Section: Collaborative Methods and Tools for Engineering and Evaluating Transparency in AI. 28 pages 9 figures, 7 tables, 1 algorithm. Source code: https://github.com/Scriptor-Group/AIMVi
CommentsUpdated title and abstract to emphasize key findings on the debiasing paradox for improved discoverability. Content and findings unchanged. 11 pages, 17 figures, Accepted at IEEE Conference on Artificial Intelligence (IEEE CAI) 2025. Full Paper acceptance in the Vertical HUMAN-CENTERED AI category
Journal ref2025 IEEE Conference on Artificial Intelligence (CAI)
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
将搜索与训练解耦:通过模型合并实现大规模语言模型预训练的数据混合缩放
Shengrui Li, Fei Zhao, Kaiyan Zhao, Jieying Ye, Haifeng Liu, Fangcheng Shi, Zheyong Xie, Yao Hu, Shaosheng Cao
机构
*
NLP Team, Xiaohongshu Inc., Shanghai, China(小红书自然语言处理团队,小红书公司,上海,中国)
;
Tsinghua University, Beijing, China(清华大学,北京,中国)
;
The University of Tokyo, Tokyo, Japan(东京大学,东京,日本)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI
AI总结
提出DeMix框架,通过模型合并预测最优数据配比,在降低搜索成本的同时提升基准性能。
Comments18 pages, 5 figures, accepted at ICML 2026
Chenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(人工智能学院,中国科学院大学,北京,中国)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
通过项目反应理论诊断LLM作为评判者的可靠性
Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, Bugeun Kim
机构
*
Department of Artificial Intelligence, Chung-Ang University, Seoul, Republic of Korea(Chung-Ang 大学人工智能系)
;
Department of Industrial Engineering, Seoul National University, Seoul, Republic of Korea(首尔国立大学工业工程系)
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
使用生成式大语言模型评估自动语音识别
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu, Mickael Rouvier, Jane Wottawa, Richard Dufour
机构
*
Idiap Research Institute(Idiap研究 institute)
;
EPFL(瑞士联邦理工学院)
;
Brno University of Technology(布拉格技术大学)
;
Avignon University(阿维尼翁大学)
;
Le Mans University(勒曼大学)
;
Nantes University(南特大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.CL
机构
*
Department of Electrical Engineering, Indian Institute of Technology Delhi(印度理工学院德里分校电子工程系)
;
Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里分校人工智能学院)
;
Hewlett Packard Enterprise, India(印度惠普企业公司)
;
Department of Computer Science & Engineering, Indian Institute of Technology Kharagpur(印度理工学院Khargapur分校计算机科学与工程系)
;
A.K.Choudhury School of Information Technology, University of Calcutta, India(印度加尔各答大学信息科技学院)
;
Department of Computer Science & Engineering, Indian Institute of Technology Bombay(印度理工学院孟买分校计算机科学与工程系)
;
Department of Computer Science, Ashoka University, India(阿什oka大学计算机科学系)
;
Department of Computer Science & Engineering, Indian Institute of Technology Jodhpur(印度理工学院朱罗普分校计算机科学与工程系)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG
机构
*
Georgetown University(乔治城大学)
;
University of Bath(巴斯大学)
;
DEVCOM U.S. Army Research Laboratory(美国陆军研究实验室)
;
University of Maryland, College Park(马里兰大学学院公园分校)
专题命中
评测与基准
:large language model(title);language model(title);LLM(abstract_cn);pretraining(abstract)
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
New York University(纽约大学)
;
Indiana University, Bloomington(印第安纳大学,布卢明顿)
;
Northeastern University(东北大学)
;
University College London(伦敦大学学院)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
ERGeoBench:多模态大语言模型中具身推理与地理定位的综合基准
Kaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou, Kaoyan Lu, Yu Feng, Yifan Zhu, Haoran Luo
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)
;
School of Materials Science and Engineering(材料科学与工程学院)
;
China Mobile Research Institute(中国移动研究院)
;
College of Computing and Data Science(计算与数据科学学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
DTBench:文档到表格提取的合成基准
Yuxiang Guo, Zhuoran Du, Nan Tang, Kezheng Tang, Congcong Ge, Yunjun Gao
机构
*
Zhejiang University(浙江大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science(香港科学与技术大学)
专题命中
评测与基准
:LLM(summary_cn,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
Understanding the Fundamental Design Decisions of Retrieval-Augmented Generation Systems
理解检索增强生成系统的基本设计决策
Shengming Zhao, Yuchen Shao, Yuheng Huang, Jiayang Song, Zhijie Wang, Chengcheng Wan, Lei Ma
机构
*
Fudan University(复旦大学)
;
East China Normal University(华东师范大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
The University of Tokyo(东京大学)
;
Macau University of Science and Technology(澳门科学理工学院)
;
Concordia University(Concordia大学)
;
University of Alberta(阿尔伯塔大学)
;
The University of Tokyo, Japan(日本东京大学)
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);prompting(abstract)