🤖 AI Summary
This study investigates whether language models inherently require trainable word embedding tables by systematically exploring whether fixed token encodings can sustain core modeling capabilities. Based on a decoder-only Transformer architecture, we compare three input interfaces—learned embeddings, canonical binary encoding, and GF(2) invertible recoding—employing a de-parameterized input projection alongside a fixed 16-bit token encoding scheme. This work is the first to disentangle architectural necessity from empirical utility in this context. Notably, at the 1.7B parameter scale, removing approximately 100 million input parameters still enables the fixed-encoding model to achieve a substantial performance of 52.4% on HellaSwag. These results demonstrate that independent, trainable word vectors are not strictly necessary for effective language modeling, thereby establishing a feasibility benchmark for parameter-free input representations.
📝 Abstract
A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.