databases

Designs, implements, and analyzes database systems and artifacts including schemas, storage layouts, query processing and optimization, indexing, transaction and concurrency control, backup/restore, replication, and data migration; builds and maintains database instances, administration tooling, and performance and reliability monitoring for structured and semi‑structured data.

databases

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.82
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Enabling Data Dependency-based Query Optimization

Jun 11, 2024
DL
Daniel Lindner
🏛️ Hasso Plattner Institute | University of Potsdam | SAP Walldorf

Traditional database query optimization relies solely on explicit schema constraints (e.g., primary and foreign keys), overlooking abundant implicit data dependencies—such as ordering or functional dependencies—that remain undetected and unexploited. Method: This paper introduces a workload-driven, lightweight dependency discovery technique that enables millisecond-scale automatic identification of such dependencies. It tightly integrates dependency-aware optimization into both the query optimizer and execution engine, supporting SQL rewriting, dependency propagation, and subquery optimization. Contribution/Results: The approach transcends conventional schema-only optimization by establishing an end-to-end dependency-aware framework. Evaluated across five mainstream DBMSs using standard benchmarks (e.g., TPC-C), it achieves up to 10% higher throughput versus primary/foreign-key–only optimization, 22% over no-dependency optimization, and overall improvements of 5%–33%. Discovery overhead is fully amortized after a single workload execution.

Automatically identifying and validating data dependencies for optimizationEnhancing query performance using non-PK/FK dependenciesIntegrating dependencies without manual declaration or SQL rewrites

Algebraic data integration*

Mar 12, 2015
PS
Patrick Schultz
🏛️ Massachusetts Institute of Technology | Conexus AI

This paper addresses the challenge of data integration across heterogeneous databases. It proposes an algebraic approach grounded in category theory and functional programming. The method models database schemas and instances as many-sorted equational theories and their initial algebras, respectively, and employs adjoint functors to enable rigorous cross-schema data migration. Innovatively, it unifies category theory, many-sorted equational logic, and functional programming paradigms; introduces a pushout-based schema mapping construction; and defines an algebraic query language—with for/where/return syntax—endowed with formal semantics. The authors implement AQL, an open-source tool supporting formally specified schema mappings, automated data migration, and verifiable query compilation. This framework constitutes the first theoretically rigorous integration of these three foundational paradigms, simultaneously ensuring mathematical precision and enhancing the automation and reliability of data integration.

Combines functional programming, category, and database theoryDevelops algebraic approach to data integrationIntroduces query language and tool (CQL) for implementation

Beyond Relations: A Case for Elevating to the Entity-Relationship Abstraction

May 06, 2025
AD
Amol Deshpande
🏛️ University of Maryland

Contemporary relational database management systems (RDBMSs) suffer from insufficient logical data independence, reducing them to passive storage layers incapable of supporting modern architectural innovation. This paper argues that the Entity-Relationship (ER) model must serve as the native abstraction layer of RDBMSs to overcome this limitation, and it provides the first systematic theoretical justification and empirical validation of the ER model’s necessity and feasibility for ensuring logical independence. Based on this insight, we design and implement ErbiumDB—a prototype system integrating metadata-driven schema management, declarative relational semantic modeling, and runtime relationship evolution. Experimental evaluation demonstrates that ER-based abstraction significantly enhances decoupling between application and storage layers, enabling flexible, semantics-aware data management. ErbiumDB establishes a novel paradigm for intelligent database architectures and delivers a rigorously validated, extensible prototype foundation for future research and development.

Addressing insufficient logical data independence in RDBMSAdvocating shift from relational to entity-relationship modelExploring innovation via prototype system ErbiumDB design

Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.

Analyzing legacy systems for modernization using knowledge graphsIdentifying logical boundaries in large, undocumented software systemsUnderstanding dependencies to assess impact of incremental changes

Conformance Testing of Relational DBMS Against SQL Specifications

Jun 13, 2024
SL
Shuang Liu
🏛️ Renmin University of China | Tianjin University | Singapore Management University | East China Normal University | University of Science and Technology of China

This work addresses the challenge of verifying relational database management systems’ (RDBMS) compliance with SQL semantics at the standard specification level. We present the first executable Prolog reference implementation grounded in the complete formal SQL semantics defined by ISO/IEC 9075, integrated with differential fuzz testing for semantic-level black-box validation. Unlike prior approaches relying solely on crash detection or meta-transformation, our method enables end-to-end verifiable modeling of SQL standard semantics. Empirical evaluation across MySQL, TiDB, SQLite, and DuckDB uncovered 19 previously unknown vulnerabilities and 11 semantic inconsistencies—each traceable to explicit violations, omissions, or ambiguities in the SQL standard. Our approach significantly enhances the decidability and interpretability of SQL implementation correctness.

Detecting bugs and inconsistencies in major RDBMS systemsFormally defining SQL semantics for reference implementationTesting RDBMS semantic conformance to SQL specifications

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic empirical analysis on the adoption and evolution of database management systems (DBMSs) in open-source projects. By examining the code history of 362 popular GitHub Java repositories, the work combines source-code heuristics, DB-Engines rankings, ORM detection, and version tracking to uncover long-term DBMS evolution patterns. The findings reveal that MySQL and PostgreSQL are the most prevalent relational DBMSs, while Redis and MongoDB exhibit stable usage among non-relational systems. HyperSQL is frequently replaced, and a “polyglot persistence” pattern—characterized by coexistence and cross-type collaboration of multiple DBMSs—is widespread. Moreover, distinct DBMSs demonstrate significantly different propensities for replacement, highlighting nuanced evolutionary dynamics in real-world software ecosystems.

Database Management SystemsDBMS adoptionDBMS migration

This work addresses the limitation of existing Text-to-SQL evaluation benchmarks, which focus narrowly on a single task and overlook critical aspects of the full database lifecycle, including design, operation, and debugging. To bridge this gap, the authors propose DBLifeBench—the first comprehensive evaluation framework encompassing five key phases: design, implementation, execution, debugging, and maintenance. Central to this framework is Progressive-Text2SQL, a novel task grounded in structured reasoning graphs that emulates human iterative problem-solving to narrow the cognitive gap between natural language and complex SQL queries. Experimental results reveal that general-purpose large language models exhibit balanced performance across phases, whereas specialized Text-to-SQL models suffer from catastrophic forgetting outside coding-centric stages. This study establishes a systematic foundation for evaluating full-stack database intelligence.

BenchmarkingDatabase LifecycleDatabase Management

Traditional relational database systems offer a monolithic set of features, yet real-world workloads often require only a subset, leading to resource redundancy and suboptimal efficiency. To address this, this work proposes an LLM-driven approach for automatically generating customized, deployable databases from natural language workload descriptions. The method employs Feature-Oriented Domain Analysis (FODA) to decompose databases into modular components and their implementation variants, constructs a dependency graph—DBGraph—augmented with cooperate edges to capture cross-module design constraints, and leverages a multi-agent architecture comprising Main, Architect, Tester, and Refining Agents to orchestrate end-to-end synthesis. Evaluated on the TPC-C benchmark (10 warehouses), the generated system achieves 130 tpmC—outperforming PostgreSQL and MySQL—while comprising only approximately 3% of their codebase and demonstrating zero failures over 60 minutes of continuous operation.

customized databasesfeature-oriented decompositionLLM-generated systems