Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes

📅 2026-05-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether large language models must rely on trainable input embedding tables. To address this, the authors propose a novel approach that entirely eliminates trainable input embeddings by employing fixed 16-dimensional binary token codes, combined with a zero-parameter dimensional expansion and an invertible affine recoding mechanism over a finite field, while retaining the standard trainable output projection. Evaluated on a 32-layer decoder model, this method achieves validation perplexity comparable to the baseline (2.36 vs. 2.44), reducing input parameters by approximately 67.1 million. A vocabulary-independent variant of the approach also performs closely (2.39). This study presents the first demonstration that large-scale language models can completely dispense with trainable input embeddings without sacrificing performance.
📝 Abstract
Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface. For a vocabulary of size $V$, exact token identity requires only $K=\lceil \log_2 V\rceil$ bits. We replace the usual trainable $V\times d_{\text{model}}$ input embedding matrix with fixed minimal binary token codes and a zero-parameter lift to model width. In our main setting, $V=65{,}536$, so $K=16$, and tokens are represented by fixed 16-dimensional binary codes tiled to $d_{\text{model}}=1024$. We also evaluate a fully table-free variant in which codes are generated from token IDs on the fly and randomly recoded by an invertible affine transform over $\mathbb{F}_2^K$. Across matched 32-layer decoder-only models trained on approximately 17B tokens and evaluated over three independent training seeds, fixed minimal codes achieve comparable held-out validation perplexity to a standard learned-input baseline while removing 67.1M trainable input parameters. The fixed-code runs have a lower mean validation perplexity in our experiments, 2.36 versus 2.44, but the observed gap is within the measured seed-to-seed variation of 4.8\%; we therefore interpret the result as evidence that the trainable input table is not necessary, rather than as a statistically resolved superiority claim. The table-free affine-recoded variant remains close at 2.39 despite a slightly shorter training run. These results show that, in this regime, a trainable input embedding table is not necessary for useful language modeling. The output projection remains standard and trainable.
Problem

Research questions and friction points this paper is trying to address.

input embedding
language models
trainable parameters
binary token codes
embedding table
Innovation

Methods, ideas, or system contributions that make the work stand out.

fixed binary token codes
trainable embedding table removal
zero-parameter lift
table-free language modeling
invertible affine recoding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Andrey Bochkov