π€ AI Summary
This study investigates the statistical properties of compression ratios for variable-to-variable (V2V) length coding under finite-sample regimes. By analytically characterizing the joint distribution of source phrase lengths and codeword lengths, the authors derive, for the first time, closed-form integral expressions for arbitrary integer-order moments of the compression ratio, along with explicit formulas for the bias and variance constants. Building on these results, they construct a highly accurate approximation to the compression ratio distribution by integrating the moment-generating function, Edgeworth expansion, and Laplaceβs method. The framework applies to Markovian sources and leverages the generalized Kraft inequality to analyze fundamental performance limits. The theoretical analysis not only cleanly decomposes the bias structures inherent in V2V, variable-to-fixed (V2F), and fixed-to-variable (F2V) coding schemes but also reveals the bias advantages of specific constructions such as Khodak codes, substantially improving the accuracy of compression ratio approximations.
π Abstract
A variable-to-variable (V2V) length code parses a source sequence into phrases of variable length and maps each phrase to a binary codeword of, generally, a different random length. After encoding $n$ phrases, the realized compression ratio $R_n=Ξ_n/Ξ£_n$ -- total codeword length over total source-symbol count -- is the finite-sample counterpart of the code's asymptotic rate $Ο$, to which it converges only as $n\to\infty$. This paper first derives exact formulas for all integer moments of $R_n$ for a given discrete memoryless source (DMS). Specifically, we obtain a closed-form formula for every moment $\E\{R_n^k\}$ as a one-dimensional integral involving only single-phrase moment generating functions of the pair $(L,\ell)$ -- the phrase length, in source symbols, and codeword length, in bits. From these moments we derive an Edgeworth approximation to the cumulative distribution function (CDF) of $R_n$ that is substantially more accurate than the central limit theorem (CLT) approximation. Using the Laplace method of integration, we also derive explicit closed-form formulas for the bias constant $C=\lim_{n\to\infty}n(\E\{R_n\}-Ο)$ and for the variance constant $\lim_{n\to\infty}n\cdot\Var\{R_n\}$. The analysis extends to Markov sources via state-indexed matrices with a redundancy formula obtained in closed form. On the coding-theoretic side, we cast V2V length codes as finite-state encoders and apply a generalized Kraft inequality for a compression-rate lower bound, and give a structural decomposition of the bias coefficient that separates cleanly across variable-to-fixed (V2F) length codes, fixed-to-variable (F2V) length codes, and V2V length codes. Applied to the Khodak code of Bugeaud, Drmota, and Szpankowski, this decomposition shows that its improved performance is reflected in its smaller bias constant.