🤖 AI Summary
To address the poor update performance of LSM-tree-based columnar storage under mixed workloads, this paper proposes a novel hybrid row-column storage engine: it maintains an in-memory incremental row store for efficient real-time updates, and upon compaction, applies fine-grained row-to-column conversion and asynchronous compression to jointly optimize update throughput and query efficiency. Key contributions include: (1) the first integrated architecture synergizing incremental row storage with columnar storage; (2) a fine-grained conversion mechanism coupled with adaptive columnar encoding; and (3) a cost-aware background resource scheduling strategy. Experimental evaluation demonstrates that, under mixed workloads, the engine achieves significantly higher update throughput than state-of-the-art columnar systems (e.g., DuckDB), while sustaining high query performance—effectively overcoming the long-standing real-time update bottleneck in columnar storage.
📝 Abstract
This study proposes a novel storage engine, SynchroStore, designed to address the inefficiency of update operations in columnar storage systems based on Log-Structured Merge Trees (LSM-Trees) under hybrid workload scenarios. While columnar storage formats demonstrate significant query performance advantages when handling large-scale datasets, traditional columnar storage systems face challenges such as high update complexity and poor real-time performance in data-intensive applications. SynchroStore introduces an incremental row storage mechanism and a fine-grained row-to-column transformation and compaction strategy, effectively balancing data update efficiency and query performance. The storage system employs an in-memory row storage structure to support efficient update operations, and the data is converted to a columnar format after freezing to support high-performance read operations. The core innovations of SynchroStore are reflected in the following aspects:(1) the organic combination of incremental row storage and columnar storage; (2) a fine-grained row-to-column transformation and compaction mechanism; (3) a cost-based scheduling strategy. These innovative features allow SynchroStore to leverage background computational resources for row-to-column transformation and compaction operations, while ensuring query performance is unaffected, thus effectively solving the update performance bottleneck of columnar storage under hybrid workloads. Experimental evaluation results show that, compared to existing columnar storage systems like DuckDB, SynchroStore exhibits significant advantages in update performance under hybrid workloads.