Value-Compressed Sparse Column (VCSC): Sparse Matrix Storage for Redundant Data

📅 2023-09-08
🏛️ Data Compression Conference
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high memory overhead and computational inefficiency of conventional sparse matrix formats (e.g., CSC, COO) when representing highly redundant sparse matrices—characterized by a small number of distinct non-zero values—this paper proposes two novel column-compressed formats: VCSC and IVCSC. VCSC integrates intra-column value deduplication and run-length counting into the CSC layout, enabling compact storage. IVCSC further enhances locality by introducing variable-length delta encoding and segmented metadata. Both formats preserve O(1) random access complexity. Evaluated on five real-world datasets, VCSC achieves an average compression ratio 1.8× higher than CSR/CSC, while IVCSC attains 2.5×, significantly outperforming state-of-the-art alternatives.
📝 Abstract
Large, sparse matrices are common in a variety of applications in scientific computing and machine learning, and are most commonly stored in CSC, CSR, or COO format. Most extensions or adaptations of these formats focus on enabling faster computation through techniques such as blocking. While common, these formats and extensions primarily take advantage of sparsity or nonzero patterns. In some cases, however, sparse data is highly redundant, with relatively few unique nonzero values. We introduce two novel extensions of CSC that aim to exploit redundancy: 1) Value-Compressed Sparse Column (VCSC) and 2) Index- and Value-Compressed Sparse Column (IVCSC). For each column, VCSC stores three arrays: unique values, value counts, and row indices, capitalizing on per-column value-redundancy by storing each unique value in a column once. IVCSC compresses further by also performing index-compression, storing an array of sections, where each section contains a value, then byte width to store row indices for that value, then positive-delta encoded row indices, then a delimiter. We summarize compression performance for VCSC and IVCSC on five varied, real-world datasets representing a wide range of use cases.
Problem

Research questions and friction points this paper is trying to address.

Handles redundant sparse data compression for efficiency
Reduces memory usage in genomics and machine learning
Extends CSC format to exploit data redundancy properties
Innovation

Methods, ideas, or system contributions that make the work stand out.

VCSC compresses redundant column values efficiently
IVCSC adds index compression via delta encoding
Both formats reduce memory usage significantly
🔎 Similar Papers
No similar papers found.
Grand Valley State University | Van Andel Institute
S
Skyler Ruiter
Grand Valley State University
S
Seth Wolfgang
Grand Valley State University
M
Marc A. Tunnell
Grand Valley State University
T
Timothy J. Triche
Van Andel Institute
E
Erin Carrier
Grand Valley State University
Z
Zachary DeBruine
Grand Valley State University, Van Andel Institute