The Discrete-Log Clock: How a Transformer Learns Modular Multiplication

📅 2026-06-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the puzzling dense spectral patterns exhibited by small Transformers on modular multiplication tasks. The authors propose replacing the conventional additive discrete Fourier transform (DFT) with a multiplicative feature transformation that aligns with the underlying multiplicative group structure. By integrating discrete logarithm-based reordering, Gini coefficient–based sparsity measurements, and MLP neuron tuning analysis, they demonstrate for the first time that the model implicitly converts multiplication into addition in the discrete logarithm domain. This approach increases the embedding spectrum’s Gini coefficient from 0.07 to 0.58, with 96.9% of MLP neurons precisely tuned to a single multiplicative frequency, thereby confirming the existence of a “discrete logarithm clock” mechanism. These findings establish a novel paradigm for neural network interpretability grounded in algebraic structure alignment.
📝 Abstract
When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies suffices. We show this density is an artifact of analyzing in the wrong basis. The natural Fourier transform for multiplication is not the standard additive DFT but the multiplicative character transform, which decomposes functions on the multiplicative group $(\mathbb{Z}/p\mathbb{Z})^*$ into its irreducible representations. Applying this transform to a grokked transformer trained on $a \cdot b \bmod 113$, we find the embedding spectrum becomes highly sparse (Gini coefficient 0.58 vs. 0.07 in the additive basis) with only 4 key frequencies carrying significant energy. Furthermore, 96.9% of MLP neurons are cleanly tuned to a single multiplicative frequency, and neuron activation heatmaps reveal 2D-periodic structure when reordered by the discrete logarithm. These results demonstrate the transformer reduces multiplication to addition in discrete-log space, implementing a "Discrete-Log Clock" algorithm analogous to Nanda et al.'s Clock algorithm for addition. The methodology generalizes: matching the analysis basis to the algebraic structure of the task reveals interpretable structure where standard tools see noise.
Problem

Research questions and friction points this paper is trying to address.

modular multiplication
Fourier transform
transformer interpretability
multiplicative characters
discrete logarithm
Innovation

Methods, ideas, or system contributions that make the work stand out.

multiplicative character transform
discrete logarithm
modular multiplication
transformer interpretability
sparse spectrum