🤖 AI Summary
This study addresses the challenge that neural network weights lack explicit grammatical structure, hindering efficient compression via variable-length pattern reuse. To overcome this, we propose WeightPE, a framework that introduces grammar size as an explicit training objective for the first time. Specifically, it embeds the Re-Pair compressor within a straight-through estimator and integrates quantization-aware training to enforce approximate weight matching during optimization. By replacing fixed codebooks with hierarchical variable-length patterns, the method minimizes grammar size under a global L2 budget. When fine-tuning ViT-B/16 and ViT-L/16 architectures on CIFAR-10, WeightPE reduces grammar volumes to 0.43× and 0.38× of their respective baselines while incurring only marginal accuracy drops of 1.9 and 1.1 percentage points, thereby achieving structured and highly efficient weight compression.
📝 Abstract
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.