Score
Operates and administers database systems—especially MongoDB—by designing, deploying, configuring and maintaining instances and clusters; this includes schema/modeling for document stores, writing and optimizing queries, managing indexes, backups and restores, replication, sharding, user access and security, and monitoring and troubleshooting performance and availability. It also covers routine operational tasks such as migrations, patching, capacity planning and automating operational workflows.
Transactional cloud-native applications (e.g., payment and booking systems) face fundamental data management challenges during cloud migration—including cross-service state consistency, persistence guarantees, application lifecycle coordination, and semantic mismatches with cloud infrastructure (e.g., messaging, containerization, elastic scaling). Current database research inadequately addresses these transactional cloud-application concerns. Method: We systematically identify and formalize these challenges, construct the first open problem map for the data management community bridging database systems and cloud-native application engineering, survey migration paradigms (microservices, Actor model, stateful stream processing), and analyze their cloud-specific adaptation bottlenecks. Contribution: We propose a theoretical framework and research roadmap for co-designing transaction semantics with cloud infrastructure, establishing foundational principles for next-generation cloud-native transaction systems.
This paper addresses version control challenges for multidimensional structured data—namely, temporal evolution, spatial collaboration, and design iteration—by introducing Operational Differencing, a novel paradigm. Methodologically, it incorporates high-level semantic operations (e.g., schema changes, refactorings) into the version model; adopts an append-only branch history with a repository-free, lightweight “copy-as-branch” architecture; and enables operational query representation and future-tense execution. Contributions include: (1) the first systematic support for automatic schema-adaptive query rewriting under schema evolution; (2) precise, fine-grained diff/merge across structural transformations; (3) resolution of four out of eight canonical schema evolution challenges; and (4) a simplified versioning experience with no explicit repository and asymptotically zero branching overhead.
This work addresses the lack of standardized, quantitative assessment of environmental impact in database systems. We propose ATLAS, the first framework to model the full lifecycle carbon footprint of analytical databases by jointly accounting for embodied carbon from hardware manufacturing and operational energy consumption. ATLAS enables systematic environmental efficiency evaluation across four representative systems—DuckDB, MonetDB, Hyper, and StarRocks—through empirical energy-efficiency measurement, geography-aware carbon intensity modeling, and an open-source benchmark suite. Key findings include: (1) architectural design induces up to 2.8× difference in runtime power consumption; (2) deployment in high-carbon-intensity grids can negate over 37% of energy-saving gains; and (3) geographic location significantly modulates relative environmental advantages among architectures. Our contribution includes the first open-source benchmark for database environmental efficiency, providing quantifiable foundations for green database design and sustainable deployment decisions.
This paper investigates inconsistency repair and consistent query answering in databases subject to both general integrity constraints and binary priority relations. We propose a symmetric-difference repair model supporting fact insertion and deletion, and for the first time extend the “optimal repair” semantics to hybrid repair settings involving priorities and arbitrary constraints. We define priority-aware Pareto-optimal repairs and rigorously prove their equivalence to foundation/justified repairs under active integrity constraints. By modeling user preferences via priority relations—and integrating complexity analysis, semantic characterization of repairs, and constraint translation techniques—we systematically establish the data complexity of repair checking and inconsistency-tolerant querying (e.g., coNP-completeness) and provide a formal foundation for repair verification. Our core contribution is a unified framework that integrates priority semantics with classical repair paradigms, enabling interpretable and verifiable preference-aware data repair.
In enterprise databases, access control policy specification and enforcement are often decoupled, leading to cumbersome, non-systematic manual auditing. To address this, we propose Intent-Based Access Control for Databases (IBAC-DB), a novel model supporting semantic policy modeling and bidirectional traceability between policies and their implementations. We introduce IBACBench—the first benchmark tailored for database access control—and DePLOI, an LLM-based system featuring a task-decomposition paradigm for NL2SQL translation. DePLOI integrates domain-customized NL2SQL, role-hierarchy-aware modeling, and hybrid synthetic data generation for evaluation. Experiments on IBACBench demonstrate that DePLOI achieves significantly higher synthesis accuracy and auditing F1-score (+10 F1) than state-of-the-art baselines, validating the feasibility and robustness of automating secure access control policy deployment.
This study addresses the lack of systematic empirical analysis on the adoption and evolution of database management systems (DBMSs) in open-source projects. By examining the code history of 362 popular GitHub Java repositories, the work combines source-code heuristics, DB-Engines rankings, ORM detection, and version tracking to uncover long-term DBMS evolution patterns. The findings reveal that MySQL and PostgreSQL are the most prevalent relational DBMSs, while Redis and MongoDB exhibit stable usage among non-relational systems. HyperSQL is frequently replaced, and a “polyglot persistence” pattern—characterized by coexistence and cross-type collaboration of multiple DBMSs—is widespread. Moreover, distinct DBMSs demonstrate significantly different propensities for replacement, highlighting nuanced evolutionary dynamics in real-world software ecosystems.
Existing relational databases lack the persistent isolation and concurrent branching capabilities required by agent workloads, hindering support for speculative modifications and nonlinear state exploration. This work presents the first systematic definition and quantitative evaluation of database branching for intelligent agents, introducing BranchBench—a benchmark combining parameterized macrobenchmarks (modeling branch-mutate-evaluate cycles) and microbenchmarks (measuring branch lifecycle overhead). We evaluate systems including Neon, DoltgreSQL, TigerBeetle, Xata, and PostgreSQL, revealing that fast-branching systems suffer 5–4000× read performance degradation as branch depth increases, while fast data operating systems exhibit 25–1500× higher latency in branch creation and switching. Our experiments demonstrate that none of the current systems can effectively scale to support representative agent workloads such as software engineering tasks, fault reproduction, or Monte Carlo Tree Search (MCTS).
This work addresses the challenges of database configuration tuning—namely high cost, constrained search space, difficulty in bottleneck diagnosis, and the need for safe execution of changes—by proposing the first agent-based tuning framework. The framework dynamically explores a cross-layer joint configuration space through interaction with both the database and operating system environments, while adhering to safety constraints during parameter adjustments. It integrates a context-aware architecture, runtime monitoring, an experience memory mechanism, and feedback-driven iterative optimization to substantially enhance tuning efficiency and performance. Experimental results on MySQL and PostgreSQL demonstrate that, compared to the strongest baseline, the proposed approach achieves an average final performance improvement of 118.1% and reduces the time required to reach optimal performance by 22.6%.
该研究通过在双范畴数据库模式中引入右伴随以实现通用量化,解决关系除法查询问题,并提供模态算子解释、一阶谓词逻辑及查询优化规则。
This work addresses the limitations of traditional database migration approaches, which are often constrained to specific source-target model pairs and struggle to support general-purpose migration in heterogeneous, multi-model environments. To overcome this, the authors propose a model-driven migration framework based on a unified data model called U-Schema. By mapping diverse data models into a common intermediate representation, the framework drastically reduces the number of required transformation pathways and enables cross-paradigm migrations. It employs traceable metadata to decouple schema transformation from data migration, thereby preserving semantic consistency while enhancing structural fidelity and query behavior equivalence. Empirical evaluation—conducted in a relational-to-document database migration scenario using both synthetic datasets and the Northwind benchmark—demonstrates the approach’s effectiveness and scalability across varying data sizes.