๐ค AI Summary
This study investigates the impact of sequence length on the performance of small language models (125M-parameter Mamba and OPT) during BabyLM pretraining. Under a fixed computational budget and 100M-token training corpus, we systematically evaluate context lengths ranging from 32 to 2048 tokens across syntactic generalization and morphological analogy reasoning tasks. Results reveal a task- and architecture-dependent optimal sequence length: short contexts (โค128 tokens) suffice for basic syntactic generalization, whereas longer contexts (โฅ512 tokens) substantially improve morphological analogy reasoning; indiscriminately increasing sequence length yields diminishing returns and harms training efficiency. To our knowledge, this is the first work within the BabyLM framework to uncover the non-monotonic utility of sequence length. Our findings provide an interpretable, reproducible empirical guide for context-length configuration in lightweight language models, enabling Pareto-optimal trade-offs between model performance and training efficiency.
๐ Abstract
Transformer language models typically operate with a fixed-length context window, which has grown in step with large-scale pretraining datasets. In the BabyLM Challenge, however, many past submissions have defaulted to using much shorter sequence lengths. We examine the impact of sequence length on BabyLM pretraining, to answer the simple question: what sequence length should we be using when training Baby LMs? Using 100M-word training data and fixed compute budgets, we compare 125M-parameter Mamba and OPT models, finding that although longer is often better, the optimal length depends on both task and architecture. Shorter sequences are sufficient for grammatical generalization tasks whereas longer contexts benefit morphological analogical reasoning tasks.