NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
NorBERTo:一个基于ModernBERT架构的葡萄牙语现代编码器模型,训练于331十亿个token语料库
机构 * Itaú Unibanco ; ICTi
专题命中 检索器与排序 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
AI总结 NorBERTo基于ModernBERT架构,利用Aurora-PT语料库训练,具备长上下文支持和高效注意力机制,在葡萄牙语语义相似性、文本蕴含和分类任务中表现优异,达到最高F1和准确率。
Comments This article has already undergone formal submission, review, acceptance, and publication in the proceedings of PROPOR 2026: Proceedings of the 17th International Conference on Computational Processing of Portuguese, Vol. 1. The published version is available in the ACL Anthology at https://aclanthology.org/2026.propor-1.18/ 11 pages, 9 tables, 2 figures
Journal ref Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1