🤖 AI Summary
Existing embedding compression methods rely on fixed coding geometries, limiting their adaptability to data distributions. This work proposes OrBIT, a framework that introduces orbital dynamics into coding geometry discovery for the first time. Specifically, OrBIT constrains a shared codebook by learning local geometric structures, compensates for quantization errors via overlap graphs, and leverages global residuals to guide sequence-level bit allocation. We theoretically establish distortion control bounds for compact graphs and demonstrate that the compiled model retains a lightweight decoder. Experimental results show that OrBIT achieves 37.9× embedding table compression on GPT-2 and over 23× compression on 7B-parameter models, significantly outperforming mainstream baselines in rate-distortion performance.
📝 Abstract
Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves $37.9\times$ compression on GPT-2 and over $23\times$ on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.