design data schemas

Design and specify formal data schemas that define entities, attributes, types, relationships and cardinalities (relational and hierarchical), message and database layouts, and accompanying metadata, provenance and lifecycle/validity constraints that make implicit assumptions explicit. Build and analyze the artifacts and processes that operationalize those schemas — schema documents and mappings, validators and generated code, simulated dataset variants, versioning and evolution strategies, and scalable data-structure engineering to ensure format compatibility and interoperability across tools and systems.

designdataschemas

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.81
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AI-assisted JSON Schema Creation and Mapping

Aug 07, 2025
FN
Felix Neubauer
🏛️ University of Stuttgart

Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.

Challenges in mapping heterogeneous data formatsDifficulty in JSON Schema creation for non-expertsLack of standardized models in many domains

This work addresses the limitations of traditional database migration approaches, which are often constrained to specific source-target model pairs and struggle to support general-purpose migration in heterogeneous, multi-model environments. To overcome this, the authors propose a model-driven migration framework based on a unified data model called U-Schema. By mapping diverse data models into a common intermediate representation, the framework drastically reduces the number of required transformation pathways and enables cross-paradigm migrations. It employs traceable metadata to decouple schema transformation from data migration, thereby preserving semantic consistency while enhancing structural fidelity and query behavior equivalence. Empirical evaluation—conducted in a relational-to-document database migration scenario using both synthetic datasets and the Northwind benchmark—demonstrates the approach’s effectiveness and scalability across varying data sizes.

database migrationheterogeneous data modelsmulti-model environments

Automatically translating natural language requirements into relational database schemas remains challenging due to reliance on domain expertise, low accuracy, and poor generalization in existing approaches. Method: This paper introduces RSchema—the first large language model (LLM)-based multi-agent framework for schema generation—featuring a novel “reflection–quality assurance” dual-role collaboration mechanism. It integrates specialized role division, cross-stage error detection, and structured correction techniques. Contribution/Results: Evaluated on the newly constructed RSchema benchmark (500+ high-quality requirement-schema pairs), our method significantly outperforms state-of-the-art LLMs and conventional methods: schema accuracy and completeness improve by 28.6% and 34.1%, respectively. RSchema achieves, for the first time, end-to-end, high-fidelity, and interpretable relational schema generation without manual intervention.

Automating relational database schema design from user requirementsEnsuring accuracy in multi-agent collaboration for schema generationOvercoming limitations of rule-based and conventional deep learning methods

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Latest Papers

What's happening recently
View more

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.

graph queriesoutput schemaproperty graphs

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.

black-box convertersconverter orchestrationdata model consistency

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

Hot Scholars

VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
NP

Nenad Petrovic

Faculty of Electronic Engineering, University of Nis
Semantic TechnologyModel-Driven Software EngineeringDomain-Specific LanguagesLLM
AK

Alois Knoll

Technische Universität München
RoboticsAISensor Data FusionAutonomous Driving