A theory of platonic representations in language models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unresolved theoretical question of why internal representations of translated sentences in multilingual models exhibit cross-lingual similarity. By positing a data hierarchy hypothesis, this work employs probabilistic context-free grammars to generate synthetic languages and derives Bayesian optimal predictors. It distinguishes geometric similarity from coordinate alignment, proposing a novel perspective that inter-layer linear component subtraction enhances representational similarity, which is subsequently validated within Transformer architectures. The findings reveal that cross-lingual similarity originates from shared abstract hierarchies, successfully explaining the intermediate-layer similarity peak and language proximity effects. Furthermore, experiments on pretrained large language models confirm that subtracting linearly predictable components effectively improves cross-lingual representational similarity.
📝 Abstract
Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained on the same data. The framework explains why cross-lingual similarity peaks in middle layers, coexists with language-specific structure, and strengthens with language proximity, model quality and data exposure. It distinguishes similarity (shared neighborhood geometry) from alignment (shared coordinates), showing that the latter occurs when code-switched data, i.e. mixed-language sentences, are abundant enough. It further predicts that subtracting from each layer the component linearly predictable from the preceding one increases cross-lingual similarity, which we confirm in pretrained LLMs.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Platonic representation hypothesis
Probabilistic context-free grammars
Belief propagation
Cross-lingual similarity
Representation alignment