🤖 AI Summary
The explosive growth in model checkpoint sizes incurs substantial storage and deployment costs, yet existing compression methods either disregard tensor structure or rely on fixed formats. This work formulates lossless tensor compression as a program synthesis problem and introduces a typed domain-specific language (DSL) that captures structural patterns—such as repeated regions and floating-point fields—through invertible operators. To guide the search for compact representations, the approach incorporates checkpoint-specific production priors into a bounded A* search, automatically generating bit-accurate, self-contained compression programs. Evaluated on ten public model checkpoints totaling 2.13 TB, the method achieves a compression ratio of 33.93%, outperforming zstd, gzip, ZipNN, and DFloat11, with compression and decompression speeds of 3.60 GB/s and 6.61 GB/s, respectively.
📝 Abstract
Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compressors rely on fixed and format-specific pipelines. We present Brevis, which formulates lossless tensor compression as program synthesis. We design a typed domain-specific language (DSL) that captures recurring tensor structures, such as repeated regions and floating-point fields, through a set of reversible operators. Given a tensor, Brevis synthesizes a self-contained DSL program that reconstructs it bit-exactly. A checkpoint-specific production prior, learned from a small representative sample of tensors, guides a bounded A* search to synthesize compact programs, which can later be executed directly for bit-exact decompression. On 10 public checkpoints spanning language, audio, and image generation models, Brevis reduces 2.13 TB of checkpoint data to 1.41 TB, a 33.93% storage reduction. It produces archives up to 30.87% smaller than those of four general-purpose compressors, including zstd and gzip, and smaller archives than the tensor-specific compressors ZipNN and DFloat11. Under a practical concurrency configuration, Brevis achieves 3.60 GB/s compression and 6.61 GB/s decompression while preserving every source byte.