Learning to Learn a Language

๐Ÿ“… 2026-10-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the heavy reliance of language models on authentic corpora, which limits their ability to infer linguistic regularities from context when encountering unseen vocabulary. To overcome this, we propose the Prior-Fitted Language Model (PF-LM), introducing a novel synthetic prior training paradigm grounded in structural causal models. Specifically, a recursive causal model generates dynamically distributed data to pretrain a 300-million-parameter byte-level Transformer, enabling next-token prediction based solely on prefixes. Without exposure to real-world vocabulary, PF-LM acquires statistical signatures and long-range dependencies, demonstrating emergent meta-learning capabilities. Evaluated across six Wikipedia languages, the model significantly reduces bits-per-byte to 0.9โ€“2.4, masters counting and approximate addition, and achieves compression performance surpassing gzip and PPMd.
๐Ÿ“ Abstract
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.
Problem

Research questions and friction points this paper is trying to address.

in-context learning
meta-learning
language modeling
synthetic pretraining
sequence prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prior-Fitted Language Model
meta-learning
structural causal model
in-context learning
synthetic pretraining
๐Ÿ”Ž Similar Papers
No similar papers found.