🤖 AI Summary
To address the high memory overhead and computational inefficiency of conventional sparse matrix formats (e.g., CSC, COO) when representing highly redundant sparse matrices—characterized by a small number of distinct non-zero values—this paper proposes two novel column-compressed formats: VCSC and IVCSC. VCSC integrates intra-column value deduplication and run-length counting into the CSC layout, enabling compact storage. IVCSC further enhances locality by introducing variable-length delta encoding and segmented metadata. Both formats preserve O(1) random access complexity. Evaluated on five real-world datasets, VCSC achieves an average compression ratio 1.8× higher than CSR/CSC, while IVCSC attains 2.5×, significantly outperforming state-of-the-art alternatives.
📝 Abstract
Large, sparse matrices are common in a variety of applications in scientific computing and machine learning, and are most commonly stored in CSC, CSR, or COO format. Most extensions or adaptations of these formats focus on enabling faster computation through techniques such as blocking. While common, these formats and extensions primarily take advantage of sparsity or nonzero patterns. In some cases, however, sparse data is highly redundant, with relatively few unique nonzero values. We introduce two novel extensions of CSC that aim to exploit redundancy: 1) Value-Compressed Sparse Column (VCSC) and 2) Index- and Value-Compressed Sparse Column (IVCSC). For each column, VCSC stores three arrays: unique values, value counts, and row indices, capitalizing on per-column value-redundancy by storing each unique value in a column once. IVCSC compresses further by also performing index-compression, storing an array of sections, where each section contains a value, then byte width to store row indices for that value, then positive-delta encoded row indices, then a delimiter. We summarize compression performance for VCSC and IVCSC on five varied, real-world datasets representing a wide range of use cases.