Score
Designs, deploys, configures, operates, and automates relational (SQL) and NoSQL database systems and clusters; implements schemas, replication/partitioning, backup/restore, access control, monitoring, capacity planning, and high-availability. Analyzes and tunes query performance, storage usage, consistency/availability tradeoffs, migrations, and failure recovery procedures to ensure reliable, secure, and performant data storage and access.
Document-oriented NoSQL databases, which typically adopt eventual consistency models, struggle to support highly reliable transaction processing. This work proposes a four-phase transaction management framework that achieves conflict-serializable consistency while preserving system scalability. By integrating transaction lifecycle management, operation classification, pre-execution conflict detection, and an adaptive locking strategy, the approach effectively eliminates deadlocks and significantly reduces both transaction abort rates and latency variance. Experimental results demonstrate that the abort rate decreases from 8.3% to 4.7%, and latency variance is reduced by 34.2%. Under high concurrency, throughput improves by 6.3%–18.4%, with a 15.2% increase observed in a 9-node cluster, accompanied by a 53% reduction in abort rate.
This study addresses the challenge of enforcing data constraints in SQL Server applications by proposing an event-driven approach based on Visual Basic for Applications (VBA). The core innovation lies in the introduction of a novel "constraint-driven design" pseudocode algorithm, which transforms the rigorous enforcement of database constraints into a standardized development workflow. Notably, this algorithm exhibits cross-platform generality, enabling seamless compatibility with NoSQL databases and heterogeneous environments while significantly reducing system adaptation costs. Experimental evaluations demonstrate that the proposed method maintains optimal data quality while validating its versatility and efficiency across diverse platforms. Ultimately, this work provides a cost-effective solution for multi-platform data constraint management, bridging the gap between strict relational integrity requirements and flexible, modern deployment architectures.
This work addresses the scarcity of real-world SQL workloads and limitations of existing synthetic query generation methods—such as rigid templates, schema hallucination, and simplistic join structures—that hinder effective training of query optimizers. The authors propose SynQL, a rule-based, deterministic SQL synthesis framework that constructs abstract syntax trees by traversing foreign key graphs in database schemas, eschewing probabilistic generation to guarantee syntactic and schema-level validity. SynQL introduces a configuration vector to explicitly control join topology, analytical complexity, and predicate selectivity, while modeling core query components including multi-table joins, projections, aggregations, and range predicates. Workloads generated on TPC-H and IMDb achieve a topological entropy of 1.53 bits; tree-based models trained on this data attain R² ≥ 0.79 on test sets with sub-millisecond inference latency.
This study addresses the lack of publicly available NoSQL workloads, which has hindered research on optimizing resource efficiency and reliability in cloud databases. To bridge this gap, the authors release the first open-source NoSQL workload derived from a real-world Cosmos DB cluster and introduce the Distressed Resource Volume (DRV) metric to quantify service quality. They further develop the LoadStar simulation framework and integrate non-parametric statistical QoS modeling, the Luna load forecasting model, and the Orbit resource scheduling algorithm to optimize replica placement and rebalancing. Experimental results demonstrate that Orbit supports higher workloads with lower error rates and reduces resource consumption by up to 35%. Already deployed in production, this approach yields an estimated annual savings exceeding $100 million while significantly enhancing service reliability for millions of users.
In enterprise databases, access control policy specification and enforcement are often decoupled, leading to cumbersome, non-systematic manual auditing. To address this, we propose Intent-Based Access Control for Databases (IBAC-DB), a novel model supporting semantic policy modeling and bidirectional traceability between policies and their implementations. We introduce IBACBench—the first benchmark tailored for database access control—and DePLOI, an LLM-based system featuring a task-decomposition paradigm for NL2SQL translation. DePLOI integrates domain-customized NL2SQL, role-hierarchy-aware modeling, and hybrid synthetic data generation for evaluation. Experiments on IBACBench demonstrate that DePLOI achieves significantly higher synthesis accuracy and auditing F1-score (+10 F1) than state-of-the-art baselines, validating the feasibility and robustness of automating secure access control policy deployment.
Existing relational databases lack the persistent isolation and concurrent branching capabilities required by agent workloads, hindering support for speculative modifications and nonlinear state exploration. This work presents the first systematic definition and quantitative evaluation of database branching for intelligent agents, introducing BranchBench—a benchmark combining parameterized macrobenchmarks (modeling branch-mutate-evaluate cycles) and microbenchmarks (measuring branch lifecycle overhead). We evaluate systems including Neon, DoltgreSQL, TigerBeetle, Xata, and PostgreSQL, revealing that fast-branching systems suffer 5–4000× read performance degradation as branch depth increases, while fast data operating systems exhibit 25–1500× higher latency in branch creation and switching. Our experiments demonstrate that none of the current systems can effectively scale to support representative agent workloads such as software engineering tasks, fault reproduction, or Monte Carlo Tree Search (MCTS).
This work addresses the challenges of Text-to-SQL in large analytical databases, where complex schemas, ambiguous parsing, and data-dependent decisions hinder performance, and conventional fixed-pipeline systems struggle to recover from early errors. To overcome these limitations, the authors propose FlexSQL, an agent-based framework that integrates dynamic schema retrieval and a flexible execution mechanism, enabling exploration of schema structures, data validation, and backtracking for correction at any reasoning stage. FlexSQL supports dual-level repair—both at the code and planning levels—and combines SQL/Python hybrid generation with multi-interpretation execution plans. Evaluated on the Spider2-Snow benchmark, this approach achieves 65.4% accuracy, outperforming stronger open-source baselines and delivering over a 10% performance gain when integrated into general-purpose programming agents.
Traditional relational database systems offer a monolithic set of features, yet real-world workloads often require only a subset, leading to resource redundancy and suboptimal efficiency. To address this, this work proposes an LLM-driven approach for automatically generating customized, deployable databases from natural language workload descriptions. The method employs Feature-Oriented Domain Analysis (FODA) to decompose databases into modular components and their implementation variants, constructs a dependency graph—DBGraph—augmented with cooperate edges to capture cross-module design constraints, and leverages a multi-agent architecture comprising Main, Architect, Tester, and Refining Agents to orchestrate end-to-end synthesis. Evaluated on the TPC-C benchmark (10 warehouses), the generated system achieves 130 tpmC—outperforming PostgreSQL and MySQL—while comprising only approximately 3% of their codebase and demonstrating zero failures over 60 minutes of continuous operation.
This work addresses the challenges of scalability, latency, and availability in multi-region, high-concurrency OLTP systems deployed in cloud environments by proposing a serverless SQL database architecture that decouples computation, storage, and transaction coordination. The system supports active-active multi-region deployments and innovatively defers coordination latency to the commit phase. By integrating stateless query processors based on Firecracker MicroVMs, a distributed arbitrator, multiversion concurrency control, and precise timestamping, it achieves strong consistency and full ACID semantics while significantly reducing cross-region latency. Experimental results demonstrate that the system can elastically scale from zero to millions of transactions per second and maintains continuous availability even in the face of region- or availability-zone-level failures.
This work addresses the challenges of database configuration tuning—namely high cost, constrained search space, difficulty in bottleneck diagnosis, and the need for safe execution of changes—by proposing the first agent-based tuning framework. The framework dynamically explores a cross-layer joint configuration space through interaction with both the database and operating system environments, while adhering to safety constraints during parameter adjustments. It integrates a context-aware architecture, runtime monitoring, an experience memory mechanism, and feedback-driven iterative optimization to substantially enhance tuning efficiency and performance. Experimental results on MySQL and PostgreSQL demonstrate that, compared to the strongest baseline, the proposed approach achieves an average final performance improvement of 118.1% and reduces the time required to reach optimal performance by 22.6%.