🤖 AI Summary
This work investigates whether large language models must rely on trainable input embedding tables. To address this, the authors propose a novel approach that entirely eliminates trainable input embeddings by employing fixed 16-dimensional binary token codes, combined with a zero-parameter dimensional expansion and an invertible affine recoding mechanism over a finite field, while retaining the standard trainable output projection. Evaluated on a 32-layer decoder model, this method achieves validation perplexity comparable to the baseline (2.36 vs. 2.44), reducing input parameters by approximately 67.1 million. A vocabulary-independent variant of the approach also performs closely (2.39). This study presents the first demonstration that large-scale language models can completely dispense with trainable input embeddings without sacrificing performance.
📝 Abstract
Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface. For a vocabulary of size $V$, exact token identity requires only $K=\lceil \log_2 V\rceil$ bits. We replace the usual trainable $V\times d_{\text{model}}$ input embedding matrix with fixed minimal binary token codes and a zero-parameter lift to model width. In our main setting, $V=65{,}536$, so $K=16$, and tokens are represented by fixed 16-dimensional binary codes tiled to $d_{\text{model}}=1024$. We also evaluate a fully table-free variant in which codes are generated from token IDs on the fly and randomly recoded by an invertible affine transform over $\mathbb{F}_2^K$. Across matched 32-layer decoder-only models trained on approximately 17B tokens and evaluated over three independent training seeds, fixed minimal codes achieve comparable held-out validation perplexity to a standard learned-input baseline while removing 67.1M trainable input parameters. The fixed-code runs have a lower mean validation perplexity in our experiments, 2.36 versus 2.44, but the observed gap is within the measured seed-to-seed variation of 4.8\%; we therefore interpret the result as evidence that the trainable input table is not necessary, rather than as a statistically resolved superiority claim. The table-free affine-recoded variant remains close at 2.39 despite a slightly shorter training run. These results show that, in this regime, a trainable input embedding table is not necessary for useful language modeling. The output projection remains standard and trainable.