Score
Design and implement methods and analyses that measure or estimate the bitrate (bits per unit) or entropy rate produced or required by a system, including estimators for compressed bitstreams, key‑stream or information entropy, and procedures to compute bits per sample. Use these estimates to compare bitrate across configurations and to analyze trade‑offs between bitrate and reconstruction quality, leakage, or other performance metrics.
Accurate estimation of entropy, mutual information, and conditional mutual information in software engineering is often hindered by high computational cost and long runtime. This paper systematically evaluates 18 bias-corrected entropy estimators across varying sample sizes and domain cardinalities. Through large-scale simulations of random joint distributions—complemented by rigorous statistical bias analysis and quantitative convergence assessment—we identify, for the first time, that the Chao–Shen and Chao–Wang–Jost estimators consistently exhibit rapid convergence and strong robustness across all entropy measures. Crucially, they achieve superior accuracy and faster convergence under low-sample-size conditions. Our findings yield a lightweight, reliable, and plug-and-play entropy estimation framework, directly applicable to software confidentiality analysis, test adequacy assessment, and machine learning feature selection.
This study addresses the problem of compressing random variables under three simultaneous constraints: distortion (fidelity), perceptual naturalness, and a newly introduced deception constraint that requires reconstructed samples to appear as if drawn from a target distribution. By incorporating this deception constraint into the classical rate–distortion framework, the work extends traditional information-theoretic analysis to cross-distribution camouflage scenarios. Leveraging an information-theoretic formulation combined with distributional distance measures, the authors derive fundamental limits of compression under deception and establish a precise trade-off among rate, distortion, and deception. This theoretical advance provides a new foundation for applications such as privacy-preserving data release and adversarial data obfuscation.
Conventional metrics such as floating-point operations per second (FLOPs) fail to capture the intrinsic performance characteristics of emerging computing paradigms—including low-precision, analog, quantum, and reversible logic—due to their hardware- and precision-specific assumptions. Method: This paper proposes a general, information-theoretic framework for computational performance evaluation, modeling computation as an information-transformation channel from input to output and using mutual information as the core metric to quantify a system’s capacity to encode, process, and preserve semantically meaningful information. Contribution/Results: It is the first work to systematically integrate Shannon’s mutual information into computational performance assessment, thereby decoupling evaluation from underlying hardware implementations and numerical representations. The framework enables paradigm-agnostic, implementation-independent performance analysis across heterogeneous computing models. It establishes a foundational theoretical basis and provides a scalable, principled metric for rigorously assessing both the effectiveness and efficiency of next-generation heterogeneous computing systems.
This work addresses three key challenges in quantization evaluation for neural codecs: high training overhead, unreliable gradient approximation, and the absence of an efficient evaluation paradigm. We propose a lightweight, simulation-driven framework for quantization effect assessment. Our method models nonlinear quantization behavior in large-scale models using low-complexity codecs and synthetically generated data with controllable bitwidths. It integrates statistical quantization simulation, soft/hard quantization annealing, and an enhanced straight-through estimator (STE) to identify and mitigate STE instability. Compared to full-scale training, our approach reduces evaluation cost by an order of magnitude in both time and hardware resources. Furthermore, it systematically uncovers the differential impact of various gradient approximation strategies on downstream performance. Extensive validation on an internal audio codec and the Descript Audio Codec demonstrates both effectiveness and cross-architecture generalizability.
This paper addresses the end-to-end Quality of Experience (QoE) assurance challenge for video streaming over best-effort networks. It systematically analyzes bottlenecks across the full pipeline—from video acquisition and compression (H.264/HEVC/AV1), upload, transcoding, CDN scheduling, adaptive bitrate (ABR) decision-making, to playback. We propose the first unified end-to-end pipeline analytical framework, classifying and modeling over 200 works along two orthogonal dimensions: methodology (heuristic, optimization, machine learning) and technical characteristics (codecs, super-resolution, etc.). The resulting methodology map is the most comprehensive to date, rigorously delineating performance boundaries and industrial deployment constraints for each approach. Our analysis identifies critical evolutionary trends—including ultra-low-latency live streaming, AI-native video coding, and edge-coordinated delivery—providing a systematic reference for both academic research and industry implementation.
This study investigates whether quantization can effectively mitigate the privacy risks posed by large language models’ verbatim reproduction of training data. Departing from conventional membership inference attacks, this work introduces verbatim extraction rate as a novel metric to systematically assess privacy leakage and evaluates the trade-off between memory retention and model performance under various quantization configurations. Using three sizes of the Pythia model family, two quantization algorithms, five precision levels (including 4-bit), and two evaluation corpora, the authors concurrently measure perplexity and verbatim extraction rate. Results show that while quantization reduces memorization more rapidly than it degrades performance, 4-bit quantization in the largest model still retains substantial amounts of training data, highlighting the limitations of quantization as a privacy-preserving technique and revealing its selective forgetting behavior toward memorized content.
This study presents the first empirical evaluation of entropy-conserving binarization (ECB) within a real-world CABAC framework, assessing both compression efficiency and computational overhead. Using a custom M-coder-based CABAC encoder, the authors integrated ECB alongside UEG, single-context Huffman, and HuffmanPos, conducting 2,480 bit-exact round-trip tests across synthetic data, procedurally generated images, and the Kodak dataset. Results demonstrate that ECB consistently outperforms single-context Huffman across all quantization parameters, achieving rate savings of 0.031–0.113 bits per symbol, while HuffmanPos surpasses other methods in 12 out of 15 source units. The primary driver of rate differences is attributed to context assignment rather than binarization length. To address ECB’s 7–10× higher decoding latency, the work proposes a single-pass interleaved decoding scheme to mitigate delay.