机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Sichuan University(四川大学)
;
Beijing Zhongguancun Academy(北京中关村科学院)
;
Zhejiang University(浙江大学)
;
Beijing Key Laboratory of Brain-Inspired General Intelligence Large Model(北京脑科学与类脑研究中心通用人工智能大模型北京市重点实验室)
;
Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑机智能技术重点实验室)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
如何合成高质量的预训练数据?提示设计、生成模型和源数据的系统研究
Joel Niklaus, Atsuki Yamaguchi, Michal Štefánik, Guilherme Penedo, Hynek Kydlíček, Elie Bakouch, Lewis Tunstall, Edward Emanuel Beeching, Thibaud Frere, Colin Raffel, Leandro von Werra, Thomas Wolf
机构
*
Hugging Face
;
University of Sheffield(谢菲尔德大学)
;
National Institute of Informatics(日本信息处理研究所)
专题命中
预训练与数据
:pretraining(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
机构
*
Harbin Institute of Technology, Shenzhen, School of Computer Science and Technology(哈尔滨工业大学(深圳)计算机科学与技术学院)
;
National University of Singapore, Department of Electronic and Computer Engineering(新加坡国立大学电子与计算机工程系)
专题命中
预训练与数据
:pretraining(title,abstract);LLM(abstract_cn);large language model(abstract);language model(abstract)
From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement
从大型语言模型谓词到逻辑张量网络:神经符号学在受规制采购中的投标验证
Cedric Haufe, Frieder Stolzenburg
机构
*
Harz University of Applied Sciences(哈尔茨应用科学大学)
;
Merseburg University of Applied Sciences(梅泽堡应用科学大学)
;
University of New South Wales (UNSW)(新南威尔士大学)
专题命中
预训练与数据
:language model(title,abstract);large language model(title);分类 cs.AI
Comments17 pages, 2 figures, 4 tables, extended version, with appendix
Journal refIn Diedrich Wolter and Gesina Schwalbe, editors, KI 2026: Advances in Artificial Intelligence -- 49th German Conference on AI, LNAI 16830, pages 236--243, Bremen, Germany, 2026
Towards a General Intelligence and Interface for Wearable Health Data
迈向可穿戴健康数据的通用智能与接口
Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
机构
*
Google Research(谷歌研究)
;
Google DeepMind(谷歌DeepMind)
;
University of Washington(华盛顿大学)
;
University of Oregon(俄勒冈大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
寻找隐式推理的最小参数预算:一种基于数据复杂度的语言模型缩放定律
Xinyi Wang, Shawn Tan, Shenbo Xu, Mingyu Jin, William Yang Wang, Rameswar Panda, Yikang Shen
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
University of Cambridge(剑桥大学)
;
University of Washington(华盛顿大学)
;
University of Toronto(多伦多大学)
;
University of Tokyo(东京大学)
专题命中
预训练与数据
:language model(title,abstract);large language model(abstract);pretraining(abstract);分类 cs.CL、cs.AI
机构
*
Hong Kong Polytechnic University(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
专题命中
预训练与数据
:large language model(title);language model(title);分类 cs.CL、cs.AI、cs.LG
CommentsV1.1 appeared in NeurIPS 2025 main conference; V2 adds GDN experiments, tightens others for a stronger, fairer comparison, and reorganizes sections; V3 adds Result 2.1 and Section 5.2 on how Canon layers improve hierarchical feature learning, from our Jan 2026 talk
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
先复制,后翻译:在多语言预训练中解读翻译动态
Felicia Körner, Maria Matveev, Florian Eichin, Gitta Kutyniok, Barbara Plank, Michael A. Hedderich
机构
*
MaiNLP, Center for Information and Language Processing, LMU Munich(MaiNLP,信息与语言处理中心,慕尼黑大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Department of Mathematics, LMU Munich(数学系,慕尼黑大学)
;
Department of Physics and Technology, University of Tromsø(物理与技术系,特罗姆瑟大学)
;
DLR-German Aerospace Center(德国航空航天中心)
专题命中
预训练与数据
:pretraining(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
Provable Training Data Identification for Large Language Models
大语言模型训练数据的可证明识别
Zhenlong Liu, Hao Zeng, Weiran Huang, Hongxin Wei
机构
*
Department of Statistics and Data Science, Southern University of Science and Technology(统计与数据科学系,南方科技大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
School of Computer Science, Shanghai Jiao Tong University(计算机科学学院,上海交通大学)
专题命中
预训练与数据
:large language model(title);language model(title);分类 cs.AI、cs.LG