Statistics of the Compression Ratio of a Variable-to-Variable Code: Exact Moments and Asymptotic Behavior

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the statistical properties of compression ratios for variable-to-variable (V2V) length coding under finite-sample regimes. By analytically characterizing the joint distribution of source phrase lengths and codeword lengths, the authors derive, for the first time, closed-form integral expressions for arbitrary integer-order moments of the compression ratio, along with explicit formulas for the bias and variance constants. Building on these results, they construct a highly accurate approximation to the compression ratio distribution by integrating the moment-generating function, Edgeworth expansion, and Laplace’s method. The framework applies to Markovian sources and leverages the generalized Kraft inequality to analyze fundamental performance limits. The theoretical analysis not only cleanly decomposes the bias structures inherent in V2V, variable-to-fixed (V2F), and fixed-to-variable (F2V) coding schemes but also reveals the bias advantages of specific constructions such as Khodak codes, substantially improving the accuracy of compression ratio approximations.
πŸ“ Abstract
A variable-to-variable (V2V) length code parses a source sequence into phrases of variable length and maps each phrase to a binary codeword of, generally, a different random length. After encoding $n$ phrases, the realized compression ratio $R_n=Ξ›_n/Ξ£_n$ -- total codeword length over total source-symbol count -- is the finite-sample counterpart of the code's asymptotic rate $ρ$, to which it converges only as $n\to\infty$. This paper first derives exact formulas for all integer moments of $R_n$ for a given discrete memoryless source (DMS). Specifically, we obtain a closed-form formula for every moment $\E\{R_n^k\}$ as a one-dimensional integral involving only single-phrase moment generating functions of the pair $(L,\ell)$ -- the phrase length, in source symbols, and codeword length, in bits. From these moments we derive an Edgeworth approximation to the cumulative distribution function (CDF) of $R_n$ that is substantially more accurate than the central limit theorem (CLT) approximation. Using the Laplace method of integration, we also derive explicit closed-form formulas for the bias constant $C=\lim_{n\to\infty}n(\E\{R_n\}-ρ)$ and for the variance constant $\lim_{n\to\infty}n\cdot\Var\{R_n\}$. The analysis extends to Markov sources via state-indexed matrices with a redundancy formula obtained in closed form. On the coding-theoretic side, we cast V2V length codes as finite-state encoders and apply a generalized Kraft inequality for a compression-rate lower bound, and give a structural decomposition of the bias coefficient that separates cleanly across variable-to-fixed (V2F) length codes, fixed-to-variable (F2V) length codes, and V2V length codes. Applied to the Khodak code of Bugeaud, Drmota, and Szpankowski, this decomposition shows that its improved performance is reflected in its smaller bias constant.
Problem

Research questions and friction points this paper is trying to address.

variable-to-variable code
compression ratio
exact moments
asymptotic behavior
finite-sample statistics
πŸ”Ž Similar Papers
No similar papers found.