What is the Best Sequence Length for BABYLM?

๐Ÿ“… 2025-10-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates the impact of sequence length on the performance of small language models (125M-parameter Mamba and OPT) during BabyLM pretraining. Under a fixed computational budget and 100M-token training corpus, we systematically evaluate context lengths ranging from 32 to 2048 tokens across syntactic generalization and morphological analogy reasoning tasks. Results reveal a task- and architecture-dependent optimal sequence length: short contexts (โ‰ค128 tokens) suffice for basic syntactic generalization, whereas longer contexts (โ‰ฅ512 tokens) substantially improve morphological analogy reasoning; indiscriminately increasing sequence length yields diminishing returns and harms training efficiency. To our knowledge, this is the first work within the BabyLM framework to uncover the non-monotonic utility of sequence length. Our findings provide an interpretable, reproducible empirical guide for context-length configuration in lightweight language models, enabling Pareto-optimal trade-offs between model performance and training efficiency.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
๐Ÿ“ Abstract
Transformer language models typically operate with a fixed-length context window, which has grown in step with large-scale pretraining datasets. In the BabyLM Challenge, however, many past submissions have defaulted to using much shorter sequence lengths. We examine the impact of sequence length on BabyLM pretraining, to answer the simple question: what sequence length should we be using when training Baby LMs? Using 100M-word training data and fixed compute budgets, we compare 125M-parameter Mamba and OPT models, finding that although longer is often better, the optimal length depends on both task and architecture. Shorter sequences are sufficient for grammatical generalization tasks whereas longer contexts benefit morphological analogical reasoning tasks.
Problem

Research questions and friction points this paper is trying to address.

Determining optimal sequence length for BabyLM pretraining efficiency
Comparing Mamba and OPT models under fixed compute constraints
Analyzing sequence length impact on grammatical vs morphological tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluated optimal sequence length for BabyLM training
Compared Mamba and OPT models under fixed compute
Found task-dependent sequence length optimization strategy
๐Ÿ”Ž Similar Papers
No similar papers found.