Score
Designs, builds, and analyzes methods to encode, quantize, transform, and transmit latent representations (embeddings or feature maps) under explicit bitrate constraints, including lightweight adapters or mappings that support multiple bitrates and that map between latent distributions (e.g., clean↔noisy). Ensures the compression pipeline is physics-aware or physics-preserving by incorporating constraints, invariants, or loss terms so that decompressed latents produce reconstructions consistent with specified physical or domain properties.
This work addresses the lack of theoretical grounding in the trade-off between fidelity and preservation of physical observables in physics-informed compression. The authors propose an analytical framework based on the local geometry of latent spaces, revealing that anisotropic distributions of compression noise induce an inherent conflict between physical alignment loss and conventional distortion metrics. Key contributions include a rate–distortion theory formulated on tangent spaces, an alignment diagnostic tool grounded in principal eigensubspace overlap, and a geometric alignment strategy integrating latent-space sensitivity analysis, entropy modeling, and physics-gradient guidance. Experiments across multiple scientific datasets demonstrate that the proposed metrics accurately predict the actual performance trade-offs between physical consistency and data-space reconstruction fidelity.
To address the train-inference inconsistency caused by non-differentiable quantization in learned image compression, this paper proposes a two-stage training paradigm: first, end-to-end pretraining using standard differentiable quantization approximations; second, fine-tuning the decoder and entropy model under strict entropy-constrained quantization (e.g., Trellis-Coded Quantization), with supervision directly from the quantized latent variables. This work is the first to embed exact entropy-constrained quantization into the fine-tuning stage, eliminating modeling biases inherent in noise injection or rounding-based approximations and effectively bridging the gap between quantization distortion and rate-distortion optimization. Experiments on Kodak and TecNick demonstrate average BD-rate reductions of 1.0%–2.0%, up to 2.2%, with zero increase in inference complexity.
Existing learned image compression methods struggle to simultaneously achieve high distortion fidelity and perceptual realism across a wide bitrate range. To address this challenge, this work proposes the Mixture of Decoder Experts (MoDE) framework, which introduces, for the first time, dual latent representations—scalar quantization (SQ) and vector quantization (VQ)—at the decoder side to separately optimize fidelity and perceptual objectives. The framework incorporates Expert-Specific Enhancement (ESE) and Cross-Expert Modulation (CEM) mechanisms to enable complementary collaboration between the two expert decoders. Operating under a unified bitstream, MoDE supports flexible decoding and consistently achieves superior rate–fidelity–perception trade-offs across low to high bitrates, thereby validating the effectiveness of the dual-stream collaborative architecture at the decoder.
This work addresses the challenge of efficiently transmitting high-dimensional features from edge devices under stringent constraints on bandwidth, latency, and energy consumption. To this end, the authors propose a trainable, bit-wise soft quantization layer that approximates discrete step functions using multiple sigmoid functions, enabling end-to-end differentiability and task-oriented lossy compression. The method allows users to specify the desired bit-width and can be seamlessly integrated as a lightweight module at the data acquisition stage of neural networks, where it is jointly optimized with downstream tasks. Experimental results across multiple datasets demonstrate that the approach achieves 5–16× compression ratios (relative to 32-bit floating-point representations) using only 2–6 bits per feature, while maintaining accuracy nearly on par with full-precision models—significantly outperforming conventional quantization baselines.
This work addresses the challenge of efficiently compressing high bit-depth depth maps by proposing a physics-aware, end-to-end compression framework. The method first losslessly maps depth maps into three-channel images and integrates multi-wavelength physics-inspired encoding with 4-bit global quantization, achieving substantial bitrate reduction while preserving near-lossless accuracy. A hybrid Transformer-CNN architecture is designed for both encoder and decoder to enhance reconstruction quality. Evaluated on the Middlebury 2014 dataset, the approach attains 99.38% depth accuracy at merely 0.307 bits per pixel (bpp), with high PSNR performance. Compared to 8-bit quantization, it reduces the bitrate by 66% at the cost of only a 0.68 dB PSNR drop, while maintaining practical computational efficiency—encoding and decoding times are 41.48 ms and 47.45 ms, respectively.
This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.
This work addresses the high computational and communication overhead of Transformer inference in cross-device deployment by introducing rate-distortion theory into intermediate representation compression—a first in this domain. The authors propose a lossy compression framework that explicitly trades off bit rate against model accuracy. Grounded in an information-theoretic perspective, they develop an analytical framework and derive PAC-style generalization bounds linking rate and entropy gap. Experimental results on language benchmarks demonstrate that the method substantially reduces communication costs, sometimes even improving accuracy, and consistently outperforms existing sophisticated baselines, thereby validating the practical relevance of the theoretical bounds.
This work addresses the challenge of achieving high-quality image compression at ultra-low bitrates (<0.05 bpp), where balancing reconstruction fidelity and model complexity remains difficult. The authors propose FlowCodec, a framework that leverages the generative priors of pretrained text-to-image models—such as Qwen-Image-2512 and FLUX.1-dev—without requiring additional conditioning signals or auxiliary networks. FlowCodec employs a two-stage mechanism combining latent variable compression with single-step streaming to enable efficient reconstruction. By introducing only 0.54% trainable parameters relative to the backbone model, it flexibly supports multi-bitrate compression. Experimental results demonstrate that the method significantly outperforms existing approaches in perceptual metrics (LPIPS and DISTS), while simultaneously achieving higher PSNR and faster encoding speed.
This work addresses the challenge of efficiently evaluating distributional fidelity between high-dimensional, large-scale synthetic and original data. The authors propose a lossless compression framework grounded in physics-informed probabilistic modeling, which leverages arithmetic coding to transform data compression into a principled measure of distributional consistency. By comparing the compressed code lengths of synthetic and original data under this model, the method yields a global, interpretable, additive, and Shannon-optimal evaluation metric. Theoretically well-founded, the approach empirically outperforms general-purpose compressors such as gzip, achieving higher compression efficiency while effectively detecting distributional discrepancies.
Existing neural image compression methods struggle to simultaneously preserve semantic fidelity and perceptual realism at ultra-low bitrates: explicit representations often lack textural detail, while implicit ones are prone to semantic drift. This work proposes the first training-free unified framework that synergistically optimizes semantics and perception by guiding a diffusion model with explicit high-order semantic information and implicitly conveying fine-grained textures via reverse channel coding. The approach innovatively integrates explicit semantic and implicit texture representations and introduces a plug-in encoder to flexibly balance the distortion-perception trade-off. Evaluated on Kodak, DIV2K, and CLIC2020, the method achieves state-of-the-art perceptual performance, surpassing DiffC by 29.92%, 19.33%, and 20.89% in DISTS BD-Rate, respectively.