Score
Designs and builds database and structured-data schema artifacts and transformations, including relational and graph schemas, metadata and labeling schemas, schema mappings, normalization, constraint enforcement, and structured output schemas. Specifies and implements processes and tools for schema versioning, evolution, mapping, enforcement, and overall schema management, including strategies for handling schema changes.
This paper addresses version control challenges for multidimensional structured data—namely, temporal evolution, spatial collaboration, and design iteration—by introducing Operational Differencing, a novel paradigm. Methodologically, it incorporates high-level semantic operations (e.g., schema changes, refactorings) into the version model; adopts an append-only branch history with a repository-free, lightweight “copy-as-branch” architecture; and enables operational query representation and future-tense execution. Contributions include: (1) the first systematic support for automatic schema-adaptive query rewriting under schema evolution; (2) precise, fine-grained diff/merge across structural transformations; (3) resolution of four out of eight canonical schema evolution challenges; and (4) a simplified versioning experience with no explicit repository and asymptotically zero branching overhead.
This study addresses the challenge of unifying attribute graph models and SQL querying within relational databases. The authors propose reinterpreting SQL foreign key semantics as reference keys, enabling natural modeling of labeled property graphs directly on standard relational table structures. They further extend SQL to support efficient graph data insertion and complex pattern matching. This approach achieves deep integration of relational and graph models within a single system without requiring an additional storage engine. Experimental results demonstrate that the proposed method effectively enables graph structure construction and advanced graph querying capabilities, significantly enhancing relational databases’ support for graph-oriented operations.
Automatically translating natural language requirements into relational database schemas remains challenging due to reliance on domain expertise, low accuracy, and poor generalization in existing approaches. Method: This paper introduces RSchema—the first large language model (LLM)-based multi-agent framework for schema generation—featuring a novel “reflection–quality assurance” dual-role collaboration mechanism. It integrates specialized role division, cross-stage error detection, and structured correction techniques. Contribution/Results: Evaluated on the newly constructed RSchema benchmark (500+ high-quality requirement-schema pairs), our method significantly outperforms state-of-the-art LLMs and conventional methods: schema accuracy and completeness improve by 28.6% and 34.1%, respectively. RSchema achieves, for the first time, end-to-end, high-fidelity, and interpretable relational schema generation without manual intervention.
This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.
Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.
This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.
This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.
This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.
Databases continuously evolve through operations such as schema changes, version updates, and data transformations; however, existing approaches typically address these functionalities in isolation, lacking a unified abstraction. This work proposes the first integrated model that unifies continuous schema evolution, version management, and data transformation within a single framework. Built upon general-purpose computational primitives, the model supports operation provenance, conditional update propagation, and change alerts, while employing a declarative mechanism to manage the co-evolution of dependent artifacts—including views and machine learning models. A prototype system implements this framework using an enhanced, parameterized Prolly Tree—a Merkle tree–inspired data structure—to construct a relational-like engine. Experimental evaluation demonstrates that the proposed approach is both feasible and offers tunable performance across diverse evolution scenarios.
This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.