Score
Designs and trains monolingual transformer-based encoder language models (e.g., RoBERTa-style) from scratch on large single-language text corpora, producing base or large pretrained checkpoints. Builds and analyzes pretraining setups (tokenization, objectives, corpora selection, optimization) and evaluates or fine-tunes the resulting encoders for downstream NLP tasks.
To address insufficient language coverage in multilingual large language model (LLM) pretraining and the difficulty of extending post-trained models to new languages, this paper proposes a low-cost intervention early in pretraining: designing a universal multilingual tokenizer that covers significantly more languages than those present in the pretraining corpus, thereby enhancing the model’s “linguistic plasticity.” This approach empirically demonstrates—for the first time—that a unified tokenizer substantially improves zero-shot and few-shot adaptation to unseen languages without degrading performance on major languages. On cross-lingual win-rate benchmarks, language adaptation gains for newly added languages improve by up to 20.2%; even for entirely novel languages absent from both the tokenizer vocabulary and pretraining data, performance increases by up to 5%. The core contribution is the identification and empirical validation that tokenizer generalizability established early in pretraining is a decisive factor governing a multilingual model’s downstream capacity for language expansion.
This work addresses the inefficiency of full-model fine-tuning in monolingual language model development for low-resource languages, which incurs high computational costs and fails to leverage modular adaptation effectively. To overcome this limitation, the authors propose an efficient transfer learning approach that integrates a target-language-specific tokenizer, freezes the corresponding embedding layer, and fine-tunes only the remaining model parameters—departing from conventional full-parameter fine-tuning paradigms. Evaluated on Scottish Gaelic, Irish, and Quechua (with only 8.5k training samples), the method consistently outperforms baseline approaches across masked language modeling, named entity recognition, and part-of-speech tagging tasks, demonstrating its effectiveness and generalizability under extremely low-resource conditions.
This work investigates the emergence mechanism of context-text copying capability in large language model (LLM) pretraining—a foundational ability underpinning in-context learning (ICL) and retrieval-augmented generation (RAG). We observe that copying emerges with a “grokking”-like pattern: it appears significantly after training loss has plateaued, then rapidly saturates. For the first time, we identify three shared characteristics between copying emergence and grokking: (i) temporal lag relative to loss reduction, (ii) independence from dataset scale, and (iii) progressive formation of induction heads—from shallow to deep layers. Leveraging Transformer-based analysis, we monitor pretraining dynamics, localize induction heads, and conduct regularization interventions. Experiments confirm that grokking-promoting techniques—particularly regularization—accelerate and strengthen copying emergence. Our findings establish an interpretable, intervention-aware training paradigm for enhancing ICL and RAG performance.
Why do large language models (LLMs) require tokenization, and why does character-level modeling lead to performance degradation in Transformers? Method: The authors construct a *k*-order Markov data source and rigorously analyze the cross-entropy of Transformers under character-level versus token-level modeling, grounding the analysis in information-theoretic modeling capacity. They establish a provable relationship between tokenization strategies and the accuracy of sequence probability estimation. Contribution/Results: Theoretically, without tokenization, Transformers collapse to modeling only unigram character distributions, failing to capture higher-order dependencies; with appropriate tokenization, learning single-step token predictions suffices to near-optimally model the source distribution. Empirically, tokenization significantly reduces cross-entropy on high-order Markov sources. This work provides the first rigorous information-theoretic and probabilistic justification that tokenization is a necessary condition for overcoming the fundamental limitations of character-level Transformer modeling.
Existing pre-trained language models rely on static subword tokenizers, leading to suboptimal multilingual efficiency and imbalanced cross-lingual performance. To address this, we propose the first dynamic tokenization framework tailored for pre-trained language models: it dynamically identifies high-frequency subword sequences in each input and merges them on-the-fly; a lightweight hypernetwork then instantaneously generates token embeddings, enabling input-adaptive subword boundary decisions. Our method integrates a BPE-inspired intra-batch merging algorithm and is compatible with both encoder (e.g., XLM-R) and decoder (e.g., Mistral-7B) architectures. Evaluated across 14 languages, it achieves over 20% average sequence length reduction for XLM-R with less than 2% performance degradation; for English decoding, it shortens sequences by 6%, accelerates inference, and significantly improves multilingual fairness.
This work addresses the rapid yet often unstructured evolution of Transformer-based language models by proposing a practical, four-dimensional evaluation framework to distinguish substantive advances from incremental improvements and to guide cross-domain deployment. The framework holistically assesses model architecture, alignment methodologies, energy-efficiency trade-offs, and domain adaptability, encompassing key techniques such as encoder/decoder variants, long-context modeling, mixture-of-experts (MoE), retrieval augmentation, instruction tuning, and preference optimization. Through empirical analysis across vertical domains—including healthcare, finance, and law—the study quantifies the trade-offs between model scale and computational cost, redefines what constitutes an “advanced” model in real-world settings, identifies critical research gaps, and offers actionable guidelines for model selection and deployment.
This work addresses the longstanding scarcity of high-quality monolingual encoders for low-resource languages such as Latvian by systematically developing and evaluating Latvian-specific pretrained language models based on RoBERTa, DeBERTaV3, and ModernBERT architectures. These models are optimized using large-scale monolingual corpora and support long-context processing. The best-performing model, lv-deberta-base (111 million parameters), substantially outperforms existing monolingual and multilingual baselines across multiple Latvian diagnostic and linguistic benchmarks, effectively bridging the gap in high-quality pretrained resources for the language. All models and evaluation materials are publicly released to support further research and development in Latvian natural language processing.
This study addresses the lack of efficient, dedicated language models for Portuguese by introducing PortBERT, a RoBERTa-based model trained from scratch using 450GB of cleaned CulturaX data. Leveraging byte-level Byte Pair Encoding (BPE) tokenization and the fairseq framework, the authors trained both base and large variants in a stable pretraining regime across a hybrid GPU/TPU infrastructure. This work presents the first systematic investigation into balancing computational efficiency and model performance in Portuguese NLP, demonstrating that PortBERT achieves competitive or superior results on the ExtraGLUE benchmark compared to existing monolingual and multilingual models. The models are publicly released on Hugging Face and fairseq, accompanied by practical metrics including training/inference wall-clock times and fine-tuning throughput.
This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.
This work addresses the challenge of optimally allocating computational resources between general pretraining and domain-specific fine-tuning for language models in multi-domain scenarios. The authors propose a scaling-law-based optimization method that trains multiple models in parallel on general corpora and leverages empirical scaling laws to accurately predict loss across varying model sizes and data volumes, enabling reliable extrapolation to larger scales. This approach dynamically determines the optimal split of compute between general pretraining and continued domain-adaptive pretraining. Experimental results demonstrate consistent and significant performance gains across diverse model scales and computational budgets on commonsense and reasoning benchmarks, marking the first achievement of efficient, cross-scale and cross-domain resource allocation for large language models.