data systems

Designs, implements, and evaluates systems and architectures that ingest, store, process, query, transfer, and serve data—including databases, data warehouses/lakes, pipelines, stream processors, query engines, and metadata/catalog services—ensuring scalability, performance, reliability, consistency, security, and operational observability. Addresses trade-offs in data models and formats, indexing, partitioning, replication, transactionality, fault tolerance, and maintainability to meet application and workload requirements.

datasystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.81
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

TreeCat: Standalone Catalog Engine for Large Data Systems

Mar 04, 2025
KO
Keonwoo Oh
🏛️ University of Maryland

Existing metadata catalogs (e.g., Hive Metastore, Iceberg, Delta Lake) face fundamental limitations in concurrent read/write throughput, strong consistency guarantees, and expressive query capabilities for hyperscale data systems. TreeCat addresses these challenges by introducing the first high-performance, purpose-built metadata catalog engine. It employs a hierarchical data model and a path-based query language; implements MVOCC—a multi-version optimistic concurrency control protocol ensuring serializable isolation; and pioneers an associative scan execution mechanism to accelerate metadata traversal and filtering. Its custom storage format is co-designed for efficient range queries and fine-grained version management. Experimental evaluation demonstrates that TreeCat delivers strict consistency under high concurrency while achieving up to 8.2× higher throughput for range queries compared to state-of-the-art systems, significantly outperforming existing solutions across key metadata management workloads.

Addressing limitations of existing catalog systems for large-scale data.Designing a specialized database engine for efficient catalog management.Ensuring strict consistency and performance in concurrent read/write operations.

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

Beyond Performance: Measuring the Environmental Impact of Analytical Databases

Apr 26, 2025
MB
Michail Bachras
🏛️ University of Toronto

This work addresses the lack of standardized, quantitative assessment of environmental impact in database systems. We propose ATLAS, the first framework to model the full lifecycle carbon footprint of analytical databases by jointly accounting for embodied carbon from hardware manufacturing and operational energy consumption. ATLAS enables systematic environmental efficiency evaluation across four representative systems—DuckDB, MonetDB, Hyper, and StarRocks—through empirical energy-efficiency measurement, geography-aware carbon intensity modeling, and an open-source benchmark suite. Key findings include: (1) architectural design induces up to 2.8× difference in runtime power consumption; (2) deployment in high-carbon-intensity grids can negate over 37% of energy-saving gains; and (3) geographic location significantly modulates relative environmental advantages among architectures. Our contribution includes the first open-source benchmark for database environmental efficiency, providing quantifiable foundations for green database design and sustainable deployment decisions.

Evaluating architectural effects on power consumption and sustainabilityMeasuring environmental impact of analytical databasesQuantifying operational and manufacturing footprint of database systems

Transactional Cloud Applications: Status Quo, Challenges, and Opportunities

Apr 23, 2025
RL
Rodrigo Laigner
🏛️ University of Copenhagen | Delft University of Technology

Transactional cloud-native applications (e.g., payment and booking systems) face fundamental data management challenges during cloud migration—including cross-service state consistency, persistence guarantees, application lifecycle coordination, and semantic mismatches with cloud infrastructure (e.g., messaging, containerization, elastic scaling). Current database research inadequately addresses these transactional cloud-application concerns. Method: We systematically identify and formalize these challenges, construct the first open problem map for the data management community bridging database systems and cloud-native application engineering, survey migration paradigms (microservices, Actor model, stateful stream processing), and analyze their cloud-specific adaptation bottlenecks. Contribution: We propose a theoretical framework and research roadmap for co-designing transaction semantics with cloud infrastructure, establishing foundational principles for next-generation cloud-native transaction systems.

Addressing distributed computing issues in cloud applicationsChallenges in migrating transactional applications to cloudEnsuring state consistency and durability in cloud

Online Marketplace: A Benchmark for Data Management in Microservices

Mar 19, 2024
RL
Rodrigo Laigner
🏛️ University of Copenhagen | Amadeus

Existing benchmarks overlook critical data management challenges in microservice architectures—including transactional consistency, cross-service querying, event-driven processing, constraint enforcement, and heterogeneous data replication. Method: We propose the first multi-dimensional evaluation framework tailored to microservice environments, grounded in a realistic e-commerce microservice model. Our benchmark suite encompasses ACID transactions, streaming event processing, cross-service consistency constraints, and hybrid replication mechanisms. Contribution/Results: We formally define the first quantifiable evaluation criteria for microservice data management, bridging a key gap in industrial-grade distributed application assessment. Empirical evaluation across mainstream data platforms identifies fundamental performance bottlenecks and architectural limitations, while delivering reproducible optimization strategies that significantly advance the design and evolution of microservice-aware data systems.

Data ManagementMicroservices ArchitectureTransactional and Query Processing

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively evaluating the trade-offs between data consistency and coordination overhead among distributed transaction patterns—such as Saga and TCC—in business logic-intensive microservice systems prior to production deployment. The authors propose a lightweight microservice simulator grounded in Domain-Driven Design (DDD), which, for the first time, integrates DDD aggregate root modeling with multiple transaction models to decouple business logic from communication and transactional infrastructure. The framework supports configurable deployment topologies and network constraints, enabling seamless transitions from centralized to fully distributed architectures while providing a deterministic verification environment. Empirical evaluation on complex multi-aggregate systems quantifies the performance, coordination overhead, and resilience of different transaction models, substantially reducing development costs and facilitating left-shifted architectural validation.

architectural simulationconsistency modelsdistributed transactions

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

This study addresses the lack of systematic guidance for enterprise software teams in choosing between monolithic and microservices architectures. The work proposes a decision-making framework that integrates technical and organizational factors, evaluating the trade-offs of each architecture across dimensions such as scalability, reliability, deployment efficiency, and organizational complexity. The assessment is grounded in system scale, business requirements, operational maturity, and long-term maintainability. Through architectural pattern analysis, a structured evaluation model, and multiple case studies, the authors develop a practical selection methodology tailored to real-world engineering contexts. This approach offers enterprises clear architectural evolution pathways and actionable guidelines aligned with their developmental stages, thereby significantly enhancing the rationality and sustainability of system design decisions.

MicroservicesMonolithic ArchitectureOrganizational Complexity