SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
SOAP、Muon及其他:推动语言模型预训练规模
Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
在人类与大语言模型合作撰写的文本中检测大语言模型生成的令牌
Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
机构
*
School of Mathematics, University of Birmingham(伯明翰大学数学学院)
;
School of Statistics and Data Science, Shanghai University of Finance and Economics(上海财经大学统计与数据科学学院)
;
Department of Statistics, The London School of Economics and Political Science(伦敦政治经济学院统计系)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Zhongguancun Academy(中关村科学院)
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
生物信息学中的生成式人工智能:模型、应用和方法进展的系统综述
Wasimul Karim, Riasad Alvi, Sayeem Been Zaman, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Md Rafi Ur Rashid, Md Rafiqul Islam, Yakub Sebastian, Sami Azam
机构
*
Department of Computer Science and Engineering, United International University(计算机科学与工程系,国际联合大学)
;
Department of Data Science and Artificial Intelligence, Monash University(数据科学与人工智能系,墨尔本大学)
;
Department of Software Engineering, Lappeenranta-Lahti University of Technology(软件工程系,拉佩兰塔-拉赫蒂技术大学)
;
Department of Computer Science and Engineering, Pennsylvania State University(计算机科学与工程系,宾夕法尼亚州立大学)
;
Faculty of Science and Technology, Charles Darwin University(科学与技术学院,查尔斯达尔文大学)
Comments10 pages, including 8 pages of main text and references plus appendices. Working paper. The held-out DocEng 2026 competition result is under organizer embargo and is not reported
SoccerSynth Field: enhancing field detection with synthetic data from virtual soccer simulator
足球合成场地:利用虚拟足球模拟器的合成数据增强场地检测
HaoBin Qin, Jiale Fang, Keisuke Fujii
机构
*
Graduate School of Informatics, Nagoya University, Nagoya, Aichi, Japan(名古屋大学信息科学研究生院)
;
RIKEN Center for Advanced Intelligence Project, Tokyo, Tokyo, Japan(理化学研究所先进情报项目中心)
DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining
DatedGPT:通过时间感知预训练防止大语言模型中的前瞻性偏差
Yutong Yan, Raphael Tang, Zhenyu Gao, Wenxi Jiang, Yao Lu
机构
*
Department of Finance, CUHK Business School, The Chinese University of Hong Kong(香港中文大学商学院金融系,香港中文大学)
;
Centre for Artificial Intelligence, University College London(伦敦大学学院人工智能中心)
专题命中
指令微调
:large language model(title,abstract);language model(title,abstract);pretraining(title);post-training(abstract)
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles