design json schemas

Design JSON Schema documents that specify the structure, data types, required and optional fields, value constraints, and validation rules for JSON data. Include reusable subschemas and $ref links, metadata and discovery-oriented properties, and optimize schema organization and naming to support validation, reuse, interoperability, and efficient discovery.

designjsonschemas

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$162K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Elimination of annotation dependencies in validation for Modern JSON Schema

Mar 14, 2025
LA
Lyes Attouche
🏛️ Université Paris-Dauphine | Sorbonne Université | Università di Pisa | Universität Passau | Università della Basilicata

Modern JSON Schema (starting from the 2019-09 draft) introduces annotation dependencies and dynamic references, leading to semantic collapse, undecidability of satisfiability checking, and exponential schema bloat. Method: We propose the first formally verified method for eliminating annotation dependencies via semantic restructuring, static keyword analysis, and equivalence-preserving transformations—rewriting modern schemas into semantically equivalent classical schemas. Contribution/Results: We rigorously prove that exponential blowup in the worst case is unavoidable during dependency elimination, and we design a practically efficient algorithm achieving sub-exponential runtime in practice. Experimental evaluation demonstrates that our algorithm significantly outperforms the theoretical lower bound, enabling reliable schema validation and seamless integration with existing toolchains. This overcomes a critical barrier hindering the migration of classical satisfiability-checking techniques to modern JSON Schema.

Addresses limitations in Modern JSON Schema's annotation dependencies.Proves exponential schema size increase when eliminating annotation dependencies.Provides an efficient algorithm for eliminating annotation-dependent keywords.

Existing approaches to JSON Schema inclusion checking struggle to balance efficiency and completeness: rule-based methods are efficient but incomplete, while instance-generation techniques are complete yet computationally expensive. This work proposes a refutation normalization technique that synergistically combines the efficiency of rule-based reasoning with the completeness of instance generation, enabling fast and reliable inclusion checking through optimized logical inference paths. Evaluated on both real-world and synthetic datasets, the proposed method significantly outperforms state-of-the-art tools, achieving theoretical completeness while substantially improving verification efficiency. The approach effectively supports complex practical applications and advances the practical applicability boundary of JSON Schema validation technologies.

completenessefficiencyinclusion checking

Blaze: Compiling JSON Schema for 10x Faster Validation

Mar 04, 2025
JC
Juan Cruz Viotti
🏛️ Sourcemeta Ltd | Rochester Institute of Technology

To address the high runtime overhead and performance bottlenecks associated with dynamic JSON Schema validation in Web APIs, this paper proposes a compilation-based validation paradigm that transforms dynamic interpretation into static compilation of native validators. We introduce the first JSON Schema compilation framework supporting preprocessing of cross-keyword complex constraints, ensuring strict specification compliance while correcting logical flaws present in mainstream tools. Leveraging abstract syntax tree analysis, pattern specialization, finite-state machine generation, and cache-aware code generation, our approach enables end-to-end compilation from schema to highly efficient validators. Experimental evaluation demonstrates an average 10× speedup in validation latency, with up to 100× acceleration in specific scenarios; compilation time remains within seconds to minutes, and runtime incurs zero interpretation overhead.

Ensures strict adherence to JSON Schema specification.Optimizes JSON Schema validation for faster processing.Reduces validation time by 10x compared to existing tools.

AI-assisted JSON Schema Creation and Mapping

Aug 07, 2025
FN
Felix Neubauer
🏛️ University of Stuttgart

Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.

Challenges in mapping heterogeneous data formatsDifficulty in JSON Schema creation for non-expertsLack of standardized models in many domains

Latest Papers

What's happening recently
View more

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.

black-box convertersconverter orchestrationdata model consistency

This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.

embedded propertiesmetadataproperty graph schemas

This work addresses the lack of semantic interoperability in structured data (e.g., JSON, YAML, CSV) within scientific workflows, which hinders consistent interpretation. The authors introduce an RDF authoring view in the MetaConfigurator editor that leverages AI-assisted generation of RML mappings to automatically transform structured data into RDF. The system supports triple editing, SPARQL querying, and knowledge graph visualization. Key innovations include the first integration of large language models for natural language-to-SPARQL translation, bidirectional synchronization between JSON-LD and RDF triples, and ontology-aware IRI auto-completion. Demonstrated on MOF synthesis experiments, the approach successfully converts JSON protocols into semantic knowledge graphs, enabling interactive exploration of relationships between experimental conditions and outcomes, thereby significantly lowering the barrier to adopting Semantic Web technologies.

JSON dataLinked Dataontology

This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.

graph queriesoutput schemaproperty graphs

Hot Scholars

EK

Emna Ksontini

University of North Calorina Wilmington ( UNCW )
Software EngineeringAI for SEInfrastucture as Code
CX

Chaowei Xiao

University of Wisconsin - Madison/NVIDIA
Trustworthy Machine LearningAdversarial Machine LearningAI SafetyRobust AI
JD

Jennifer D'Souza

TIB Leibniz Information Centre for Science and Technology
Natural Language ProcessingScientific Knowledge ExtractionLLM EvaluationScientometrics
SK

Saket Kumar

The Mathworks Inc
Software ArchitectureGenerative AILLMAgentic RAG
AE

Abul Ehtesham

Kent State University
Generative AILarge Language ModelAgentic RAGAgentic AI