Scaling limit of the Random Language Model

📅 2026-06-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the statistical behavior and phase transition mechanisms of stochastic language models in the large-scale limit. By introducing a scaling regime—where the number of syntactic symbols \(N \to \infty\), temperature approaches zero, and the rescaled variable \(x = \tilde{\varepsilon}_d \log N\) is held fixed—and combining large deviation theory with the semi-annealed approximation, the model is rigorously mapped to a random energy model endowed with nontrivial combinatorial structure. The study establishes, for the first time, the existence of a condensation phase transition at the critical point \(x_c = 1/8\) and identifies the onset of entropy reduction at \(x = 1/2\). It further elucidates the interplay among grammar size, corpus length, and temperature that governs multiple scaling laws, delineates distinct saturation, critical, and scaling regimes, explains the origin of slow convergence in the large-\(N\) limit, and provides a universal theoretical framework for understanding empirical statistical regularities in natural language and the behavior of large language models.
📝 Abstract
We develop a quantitative theory of the Random Language Model (RLM), an ensemble of stochastic context-free grammars, in a scaling limit where the number of hidden symbols $N \to \infty$ while the grammar temperature $\tildeε_d \to 0$ at fixed $x = {\tildeε}_d \log N$. In this limit, the model admits a controlled description based on a large-deviation principle over rule-usage patterns. A semi-annealed approximation maps the problem to a class of Random Energy Models with nontrivial combinatorics. We show that the RLM exhibits a condensation transition at a critical value $x_c=1/8$, below which rule usage concentrates and language statistics acquire a nontrivial dependence on corpus length. A second characteristic scale at $x=1/2$ marks the onset of entropy reduction from its maximal value. Across these regimes, we derive explicit scaling laws for the number of distinct rules, entropy, and related observables, identifying distinct scaling, saturation, and critical regimes controlled by the interplay of grammar size, corpus length, and temperature. The theory resolves previous ambiguities regarding the existence of a thermodynamic transition and explains the slow approach to the large-$N$ limit as a consequence of the dependence on $\log N$. It further provides a unified framework in which universal statistical properties of language emerge from typical realizations of generative grammars, with implications for both natural language statistics and the behavior of large language models.
Problem

Research questions and friction points this paper is trying to address.

Random Language Model
scaling limit
condensation transition
thermodynamic transition
large-deviation principle
Innovation

Methods, ideas, or system contributions that make the work stand out.

Random Language Model
scaling limit
condensation transition
large-deviation principle
Random Energy Model
🔎 Similar Papers