columnar storage design

Designs, implements, and evaluates column-oriented on-disk and in-file storage systems and formats — including layout, serialization, compression/encoding schemes, metadata, and APIs — for tabular data to optimize analytical query performance, I/O throughput, compression effectiveness, and interoperability.

columnarstoragedesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$185K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AutoComp: Automated Data Compaction for Log-Structured Tables in Data Lakes

Apr 05, 2025
AG
Anja Gruenheid
🏛️ Microsoft | LinkedIn | University of Maryland

Log-structured table formats (e.g., Delta Lake, Iceberg, Hudi) in data lakes suffer from excessive small files due to append-only writes and metadata-heavy operations, degrading query performance, increasing storage costs, and limiting system scalability. Existing compaction mechanisms lack flexibility, pursue narrow objectives, and fail to balance benefits against operational overhead. To address this, we propose Scalable Adaptive Compaction (SAC), an extensible, workload-aware, metadata-driven framework featuring dynamic threshold tuning, lightweight online evaluation, and a modular rule engine. SAC is production-deployed via the OpenHouse control plane. Evaluated on LinkedIn’s production workloads and synthetic benchmarks, SAC reduces file counts by up to 92%, improves typical query latency by 3.8×, and maintains bounded runtime overhead.

Automates compaction for log-structured tablesImproves query performance and storage efficiencyReduces small file proliferation in data lakes

SynchroStore: A Cost-Based Fine-Grained Incremental Compaction for Hybrid Workloads

Mar 24, 2025
YZ
Yinan Zhang
🏛️ East China Normal University

To address the poor update performance of LSM-tree-based columnar storage under mixed workloads, this paper proposes a novel hybrid row-column storage engine: it maintains an in-memory incremental row store for efficient real-time updates, and upon compaction, applies fine-grained row-to-column conversion and asynchronous compression to jointly optimize update throughput and query efficiency. Key contributions include: (1) the first integrated architecture synergizing incremental row storage with columnar storage; (2) a fine-grained conversion mechanism coupled with adaptive columnar encoding; and (3) a cost-aware background resource scheduling strategy. Experimental evaluation demonstrates that, under mixed workloads, the engine achieves significantly higher update throughput than state-of-the-art columnar systems (e.g., DuckDB), while sustaining high query performance—effectively overcoming the long-standing real-time update bottleneck in columnar storage.

Balances update and query performance in hybrid workloadsImproves update efficiency in columnar storage systemsIntroduces fine-grained row-to-column transformation strategy

Existing Computational Object Storage (COS) systems face three key bottlenecks when performing large-scale scientific tabular data SQL analytics in HPC environments: rigid output formats, limited operator pushdown capability, and inadequate adaptation to deep storage hierarchies. To address these, we propose COS-SQL—a near-data SQL analytics framework tailored for HPC. Our approach features: (1) flexible output format support—including Arrow columnar layout; (2) full-stage pushdown of complex operators and array expressions; and (3) dynamic execution path selection based on hierarchical storage structure. COS-SQL adopts object-level storage organization and tightly integrates with Apache Spark. Evaluated on real-world HPC workloads, it achieves up to 32.7% end-to-end performance improvement over state-of-the-art COS systems, significantly enhancing both analytical flexibility and execution efficiency.

Expands operator support for complex analytical tasks in storage systemsOptimizes execution paths across deep storage hierarchies for efficiencyOvercomes rigid output format constraints in computation-enabled object storage

M2: An Analytic System with Specialized Storage Engines for Multi-Model Workloads

Aug 04, 2025
KK
Kyoseung Koo
🏛️ Seoul National University

To address the high communication overhead of multilingual persistent operations and the suboptimal cross-model processing efficiency caused by monolithic storage engines in existing multimodel databases, this paper proposes an integrated multimodel storage engine architecture. The architecture unifies heterogeneous storage engines, each specialized and optimized for a distinct data model; introduces a multi-stage hash join algorithm to enable efficient cross-model joins; and implements unified query plan compilation and coordinated execution across models. Experimental evaluation demonstrates that the system achieves up to 188× speedup over the best-performing baseline on representative multimodel analytical workloads, while significantly improving both performance and scalability under complex, mixed-model query loads.

Handling multiple data models efficiently in analyticsOptimizing storage engines for diverse data modelsReducing communication costs in polyglot persistence systems

High-Performance DBMSs with io_uring: When and How to use it

Dec 04, 2025
MJ
Matthias Jasny
🏛️ Technische Universität Darmstadt | Technische Universität München | TigerBeetle | DFKI

This work investigates how modern database systems can achieve low-overhead, high-throughput I/O via the Linux io_uring interface. Focusing on two representative workloads—storage-intensive buffer management and network-intensive analytical processing—we propose and empirically validate systematic design principles for integrating io_uring into databases. Our methodology centers on three key mechanisms: persistent buffer registration, passthrough I/O, and a unified I/O–network programming model, rigorously analyzing their end-to-end performance impact. Innovatively, we co-design asynchronous batched I/O, zero-copy passthrough, and architecture-level coordination to establish a portable I/O optimization framework. We implement this approach in PostgreSQL and demonstrate a 14% throughput improvement under realistic workloads. To our knowledge, this is the first systematic study to empirically confirm that fine-grained, kernel-level I/O primitive optimizations yield substantial, measurable gains at the full-system level.

Determining when io_uring benefits storage and network workloadsEvaluating io_uring for efficient database I/O operationsProviding guidelines for integrating io_uring in DBMSs

Latest Papers

What's happening recently
View more

This study addresses the challenge of selecting an optimal Data Lakehouse architecture based on data type and scale by presenting the first systematic evaluation of Apache Hudi, Apache Iceberg, and Delta Lake in terms of data ingestion efficiency and storage overhead for structured and semi-structured workloads. Conducted on the Apache Spark platform, the empirical comparison employs a four-stage ETL pipeline to assess the three frameworks under realistic conditions. Experimental results demonstrate that Delta Lake achieves the fastest data loading performance, while Iceberg excels in storage compression ratio and system stability. In contrast, Hudi exhibits comparatively lower efficiency in both batch ingestion and storage utilization. These findings provide critical empirical evidence and practical guidance for informed architectural decisions in Lakehouse deployments.

analytical data systemsApache SparkData Lakehouse

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

This study addresses the performance degradation in lakehouse tables caused by small-file accumulation, a problem exacerbated by the absence of principled merge-triggering mechanisms. The authors develop an open simulation framework that generates diverse table layouts based on Apache Iceberg and extracts 17 metadata features to predict post-merge file reduction ratios with high accuracy using XGBoost (R²=0.998). Their work is the first to systematically uncover strong correlations between metadata characteristics and merge efficacy. They further propose a model-free merging strategy relying solely on a single partition-level threshold, which demonstrates remarkable generalization across heterogeneous workloads (R²=0.976). Experimental results show that this approach substantially improves metadata-intensive query performance while also revealing its subtle impact on the parallelism of full-table scan compaction.

compactionlakehousemetadata

This study addresses the challenge organizations face in selecting among the three dominant data warehouse modeling approaches—Inmon, Kimball, and Data Vault—by developing a multidimensional evaluation framework. The framework systematically compares these methodologies across key dimensions including architectural philosophy, modeling techniques, scalability, agility, query performance, and auditability. Integrating contextual factors such as organizational size, regulatory compliance requirements, and analytical maturity, the work proposes a situational selection guideline that aligns modeling choices with strategic business objectives rather than purely technical merits. By demonstrating that no single approach is universally optimal and articulating a practical, context-sensitive decision-making standard, this research fills a critical gap in practitioner-oriented guidance for data warehouse design.

data modelling methodologiesData Vaultenterprise data warehousing

Hot Scholars

MK

Marcus Kessel

University of Mannheim
Software EngineeringData Science
FN

Frank Neven

Professor of Computer Science, Hasselt University
Databases
SV

Stijn Vansummeren

Professor of Computer Science, Hasselt University
Database TheoryDatabase SystemsSemantic WebTheoretical Computer Science
SK

Sidharth Kumar

Associate Professor, University of Illinois at Chicago
HPCParallel I/OVisualization
KM

Kristopher Micinski

Syracuse University
Programming LanguagesStatic AnalysisAutomated ReasoningReverse Engineering