Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models

📅 2026-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the geometric limits of feature representations in Transformer language models, focusing on the upper bound of nearly orthogonal directions. Building upon assumptions of linear representability and superposition, the work analyzes the boundary of cosine similarity distributions in embedding matrices and introduces a tolerable orthogonality deviation ε. Centered on ε, it proposes a refined representational capacity formula that reveals exponential sensitivity of capacity to ε. The analysis further uncovers that large models tend to strengthen orthogonality constraints rather than merely expanding dimensionality. Combining cosine similarity analysis, a modified Johnson–Lindenstrauss lemma, and modeling of near-orthogonal packing efficiency, the method validates two distinct orthogonality patterns across dozens of open-source models, reducing capacity prediction error by two orders of magnitude and establishing, for the first time, a quantifiable metric for representational capacity.
📝 Abstract
Model dimension ($d_{model}$) is a fundamental hyperparameter in transformer language models, yet its role in setting the geometric limits of feature representation remains under-explored. Grounded in the Linear Representation and Superposition Hypotheses - which propose that models encode features as near-orthogonal directions in latent space - we develop a framework for estimating how many such directions a model can support. We first establish the embedding matrix as a measurable proxy for near-orthogonality constraints across the latent space: the boundary between meaningful token relationships and incidental similarity in the pairwise cosine similarity distribution gives a concrete estimate of the model's accepted deviation $\varepsilon$ from perfect orthogonality. Applying this metric across dozens of open-source models reveals two classes: models with high $\varepsilon$ whose embeddings lack near-orthogonal structure, and models with low $\varepsilon$ that maintain it. We then show that the standard Johnson-Lindenstrauss lemma greatly underestimates the packing efficiency of trained representations, and derive an adjusted capacity formula in which the number of near-orthogonal directions depends on the ratio of vectors to dimensions ($k/d$) rather than the raw count - a single modification that cuts prediction error by two orders of magnitude with no extra parameters. Combining these results, we define representational capacity as an upper bound on the number of distinguishable directions available for features and embeddings in a model's latent space. Capacity is exponentially sensitive to $\varepsilon$, and larger models favor tighter orthogonality constraints over maximizing raw capacity - a pattern compatible with several explanations (a stability-capacity trade-off, a ceiling on usable concepts, or confounds with model scale) that we leave to future work.
Problem

Research questions and friction points this paper is trying to address.

Representational Capacity
Transformer Language Models
Near-Orthogonality
Geometric Limits
Feature Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

representational capacity
near-orthogonality
Johnson-Lindenstrauss lemma
embedding geometry
feature representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alexander Guha
Arizona State University