sequence and state‑space compression

Designs and evaluates algorithms, codecs, and tools that reduce storage or transmission cost of discrete sequences and state-space representations, including methods for sequence compression, state-space compression, and streaming decompression. Work includes developing and tuning compression parameters, implementing compression/decompression pipelines, and analyzing compression ratio, latency, and fidelity across data-compression techniques.

sequenceandstate‑spacecompression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Lossless Compression of Time Series Data: A Comparative Study

Oct 08, 2025
JG
Jonas G. Matt
🏛️ ETH Zürich | ABB

To address the need for efficient lossless compression of time-series data, this paper proposes a two-stage unified framework comprising data transformation (e.g., differencing, predictive coding) followed by entropy coding (e.g., Huffman, arithmetic coding). We conduct the largest end-to-end lossless compression benchmark to date, evaluating combinations of transformation and coding methods across synthetic and diverse real-world time-series datasets. A standardized evaluation protocol and systematic ablation analysis are introduced to quantify individual component contributions and reveal the sensitivity of algorithm performance to intrinsic data characteristics. Key contributions include: (i) empirical validation that holistic pipeline design—not isolated components—dominates overall compression efficacy; (ii) identification of critical synergies among transformation and coding modules; and (iii) establishment of principled guidelines for scenario-aware algorithm selection and customized pipeline construction, bridging theoretical insight with practical deployment.

Comparing algorithm performance across diverse synthetic and real datasetsEvaluating lossless compression methods for time series dataIdentifying optimal compression pipeline configurations for specific data characteristics

This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.

access latencyanalytical operationsdata decompression

Challenges and Solutions in Selecting Optimal Lossless Data Compression Algorithms

Sep 23, 2025
MA
Md. Atiqur Rahman
🏛️ East West University | Bangladesh University of Business and Technology (BUBT)

Lossless compression algorithms face inherent trade-offs among multidimensional performance metrics—particularly compression ratio and encoding/decoding speed—posing challenges in latency-sensitive, high-fidelity applications such as medical imaging. Method: This paper introduces the first unified, quantifiable multi-objective evaluation framework for lossless compression, employing normalized weighted modeling to dynamically balance compression ratio and speed across diverse data modalities (e.g., images, text). Contribution/Results: Its key innovation lies in a standardized multi-objective scoring model that bridges the gap between theoretical metrics and practical deployment requirements. Extensive experiments demonstrate the framework’s robustness in identifying scenario-optimal compressors: learned codecs achieve superior compression ratios, while traditional algorithms retain advantages in speed-critical tasks. The framework enables principled, application-aware algorithm selection without requiring domain-specific re-engineering.

Balancing compression ratio, encoding speed, and decoding speed requirementsProviding objective comparisons for diverse applications like medical imagingSelecting optimal lossless compression algorithms with conflicting performance trade-offs

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Sep 01, 2025
DW
Daoce Wang
🏛️ Univ of Nebraska Omaha | Los Alamos National Lab | Oakland University | Oak Ridge National Lab | Indiana University | Temple University | Florida State University

Addressing the challenge of simultaneously achieving progressive decompression, random access, and high compression quality/speed in lossy scientific data compression, this paper proposes the first unified framework supporting both decompression modes. Our method employs hierarchical data partitioning and hierarchical prediction, integrated with error-bounded lossy compression and streaming-based codec design. It maintains reconstruction accuracy comparable to state-of-the-art non-streaming compressors (e.g., SZ3), while significantly improving decompression throughput—up to 6.7× faster than SZ3. To our knowledge, this is the first approach to jointly optimize high fidelity, high throughput, and flexible data access. The framework enables practical online analysis and on-demand data retrieval, bridging a critical gap between theoretical compression efficiency and real-world scientific computing workflows.

Enabling progressive and random-access decompression without quality lossMaintaining high compression speed while supporting streaming featuresOvercoming partitioning impact to achieve state-of-the-art compression quality

The field of bounded-lossy compression for scientific data lacks a systematic survey and unified classification framework. Method: This work proposes the first six-dimensional taxonomy, categorizing existing approaches into six model classes; systematically analyzes 46 state-of-the-art compressors—elucidating their design principles, error-control mechanisms, and application domains—and distills five core technical components: predictive coding, transform coding, quantization, entropy coding, and parallel/distributed architectures. Furthermore, it establishes domain-specific compression design methodologies and selection guidelines tailored to high-performance computing (HPC), climate modeling, and particle physics. Contribution/Results: The framework enables high-fidelity, high-ratio scientific data management and has become a benchmark reference across multiple disciplines.

Design compressors for specific scientific applicationsEvaluate 46 state-of-the-art lossy compressorsSurvey error-bounded lossy compression techniques

Latest Papers

What's happening recently
View more

This study addresses the optimization of general-purpose lossless compression under realistic resource constraints—specifically, memory usage capped at 8 GB and decompressor size limited to 1 MB—by organizing an international challenge based on a public training set and a hidden test set comprising 16 heterogeneous files. Performance is evaluated multidimensionally using compression ratio, compression/decompression time, Weissman score, and Pareto front analysis. The generalization capability of submissions is further assessed on external large-scale datasets, while Normalized Compression Distance (NCD) is employed to analyze inter-solution relationships. Among 117 valid submissions, several outperformed mainstream tools on external data, underscoring the critical role of advanced probabilistic modeling and effective multi-objective trade-offs in enhancing compression performance.

Algorithmic Information Theorycompression benchmarkgeneralization

General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.

application-specific compressorslossless compressionmaintainability

This study addresses the trade-off between compression efficiency and transmission delay when using variable-length coding for real-time text streaming over fixed-rate channels, where queueing delays arise from fluctuating codeword lengths. For the first time, it systematically evaluates predictive coding driven by large language models in this context, employing GPT-2 (124M) and Llama 3.2 (3B) as causal predictors combined with entropy coding schemes including Shannon, Huffman, arithmetic coding, rANS, and gzip. Experimental results demonstrate that increasing model size by a factor of 25 reduces bits per character by approximately 38%. Among entropy coders, Huffman coding proves more practical in over-provisioned channels due to its zero algorithmic delay, whereas arithmetic coding approaches the theoretical compression limit at the cost of introducing decoding latency.

causal language modelcompression-delay tradeoffentropy coding

Traditional three-stage Huffman coding incurs significant computational and latency overhead in multi-accelerator communication due to the need for real-time frequency analysis, codebook generation, and transmission, making it ill-suited for low-latency requirements. This work proposes a single-stage Huffman encoder that constructs a fixed codebook based on the average probability distribution derived from historical data batches, eliminating the need for real-time codebook generation and transmission. For the first time, this approach is applied to lossless compression of cross-layer and sharded tensors. Evaluated on the Gemma 2B model, the method achieves a compression ratio only 0.5% lower than per-shard Huffman coding and lies within 1% of the Shannon limit, substantially improving communication efficiency while closely approaching the theoretical compression bound.

Huffman codinglarge language modelslatency-sensitive communication

In industrial and IoT applications, high-ratio compression of real-time/historical process data reduces storage costs and improves efficiency but risks degrading the accuracy of statistical analysis, anomaly detection, and machine learning models. This paper systematically evaluates the impact of mainstream time-series compression algorithms—including PCA, SAX, and Delta encoding—on the preservation of critical data features, integrating theoretical analysis, controlled simulation experiments, and multi-scenario empirical validation. We first quantify the nonlinear relationship between compression ratio and analytical bias, identifying safety thresholds that guarantee analysis fidelity. Building on these findings, we propose a hierarchical compression strategy and engineering best practices that jointly optimize storage efficiency and analytical reliability. Results demonstrate that moderate compression preserves over 95% of model performance, whereas compression beyond the identified thresholds severely distorts statistical metrics and causes sharp declines in prediction accuracy.

Assessing trade-offs between compression efficiency and analytical reliabilityEvaluating how data compression affects analytical solution accuracyIdentifying optimal compression methods to preserve critical data patterns

Hot Scholars

SD

Sheng Di

Argonne National Labratory, IEEE Senior Member
HPCData CompressionResilienceCloud/Grid Computing/P2P
FC

Franck Cappello

Argonne National Laboratory, IEEE Fellow
Parallel ProcessingParallel ComputingHigh Performance ComputingFault Tolerance
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
JK

Jan Kautz

Vice President of Research, NVIDIA Research
Computer VisionMachine LearningVisual Computing
BZ

Bo Zheng

Researcher, Alibaba Group
AINetworkE-Commerce