Emergent phases of superposition: from partial to full representation

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates how the number of features, data distribution, and model width jointly determine representational configurations and loss within the superposition mechanism of large language models. Building upon Anthropic’s toy model and employing partial random projection approximations, we conduct theoretical derivations corroborated by experimental validation. Our analysis reveals that as model width increases, the system undergoes a continuous phase transition from partial to full representation. We propose a theory demonstrating the linear scaling of critical width, quantify the triggering effect of non-uniform data distributions, and characterize the evolution of loss across this phase transition. This work elucidates the intrinsic mechanisms through which model width and data statistics collaboratively shape representations, offering a novel perspective for understanding representational scaling in large models.
πŸ“ Abstract
Large language models are thought to represent features by vectors in a hidden space of dimension given by the model's width. Superposition, in which more features are represented than the width by letting representation vectors overlap, is a leading account of how representation vectors are organized. However, how model width and data statistics determine the configuration of representation vectors and the resulting loss when the number of features and the width are large remains less understood. Here we show, in Anthropic's toy model of superposition, that increasing the width drives a continuous phase transition from a partial-representation phase, where only a subset of features receives appreciable representation vectors while the rest vanish, to a full-representation phase, where every feature is represented. Our theory via a partial random projection approximation predicts, and experiments confirm, that the critical width grows linearly with the number of active features up to a logarithmic factor. The loss scaling changes across the transition: below the critical width, the loss grows linearly with the number of active features and depends weakly on the width in a form set by data statistics; above it, the loss grows approximately quadratically with the number of active features and decays inversely with the width. Non-uniform firing probabilities delay the transition and lower the loss, as more frequent features occupy more space. Our results provide an account of how model width and data statistics jointly shape representations and loss, a step toward understanding representation scaling in large models.
Problem

Research questions and friction points this paper is trying to address.

superposition
representation
model width
phase transition
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

superposition
phase transition
representation scaling
random projection approximation
loss scaling
πŸ”Ž Similar Papers
No similar papers found.