PortBERT: Navigating the Depths of Portuguese Language Models

📅 2026-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of efficient, dedicated language models for Portuguese by introducing PortBERT, a RoBERTa-based model trained from scratch using 450GB of cleaned CulturaX data. Leveraging byte-level Byte Pair Encoding (BPE) tokenization and the fairseq framework, the authors trained both base and large variants in a stable pretraining regime across a hybrid GPU/TPU infrastructure. This work presents the first systematic investigation into balancing computational efficiency and model performance in Portuguese NLP, demonstrating that PortBERT achieves competitive or superior results on the ExtraGLUE benchmark compared to existing monolingual and multilingual models. The models are publicly released on Hugging Face and fairseq, accompanied by practical metrics including training/inference wall-clock times and fine-tuning throughput.
📝 Abstract
Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neglecting training and deployment efficiency. In the present work, we introduce PortBERT, a family of RoBERTa-based language models for Portuguese, designed to balance performance and efficiency. Trained from scratch on over 450 GB of deduplicated and filtered mC4 and OSCAR23 from CulturaX using fairseq, PortBERT leverages byte-level BPE tokenization and stable pre-training routines across both GPU and TPU processors. We release two variants, PortBERT base and PortBERT large, and evaluate them on ExtraGLUE, a suite of translated GLUE and SuperGLUE tasks. Both models perform competitively, matching or surpassing existing monolingual and multilingual models. Beyond accuracy, we report training and inference times as well as fine-tuning throughput, providing practical insights into model efficiency. PortBERT thus complements prior work by addressing the underexplored dimension of compute-performance tradeoffs in Portuguese NLP. We release all models on Huggingface and provide fairseq checkpoints to support further research and applications.
Problem

Research questions and friction points this paper is trying to address.

Portuguese language models
model efficiency
compute-performance tradeoffs
NLP
language-specific models
Innovation

Methods, ideas, or system contributions that make the work stand out.

PortBERT
language-specific efficiency
byte-level BPE
compute-performance tradeoff
Portuguese NLP
🔎 Similar Papers
No similar papers found.
R
Raphael Scheible-Schmitt
School of Computation, Information and Technology, Technical University of Munich, Munich, Germany
H
Henry He
School of Computation, Information and Technology, Technical University of Munich, Munich, Germany
A
Armando B. Mendes
IS2E - Intelligent Systems, Science and Engineering, LIACC polo on Azores University, Ponta Delgada, Portugal