Score
Design JSON Schema documents that specify the structure, data types, required and optional fields, value constraints, and validation rules for JSON data. Include reusable subschemas and $ref links, metadata and discovery-oriented properties, and optimize schema organization and naming to support validation, reuse, interoperability, and efficient discovery.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
Modern JSON Schema (starting from the 2019-09 draft) introduces annotation dependencies and dynamic references, leading to semantic collapse, undecidability of satisfiability checking, and exponential schema bloat. Method: We propose the first formally verified method for eliminating annotation dependencies via semantic restructuring, static keyword analysis, and equivalence-preserving transformations—rewriting modern schemas into semantically equivalent classical schemas. Contribution/Results: We rigorously prove that exponential blowup in the worst case is unavoidable during dependency elimination, and we design a practically efficient algorithm achieving sub-exponential runtime in practice. Experimental evaluation demonstrates that our algorithm significantly outperforms the theoretical lower bound, enabling reliable schema validation and seamless integration with existing toolchains. This overcomes a critical barrier hindering the migration of classical satisfiability-checking techniques to modern JSON Schema.
Existing approaches to JSON Schema inclusion checking struggle to balance efficiency and completeness: rule-based methods are efficient but incomplete, while instance-generation techniques are complete yet computationally expensive. This work proposes a refutation normalization technique that synergistically combines the efficiency of rule-based reasoning with the completeness of instance generation, enabling fast and reliable inclusion checking through optimized logical inference paths. Evaluated on both real-world and synthetic datasets, the proposed method significantly outperforms state-of-the-art tools, achieving theoretical completeness while substantially improving verification efficiency. The approach effectively supports complex practical applications and advances the practical applicability boundary of JSON Schema validation technologies.
To address the high runtime overhead and performance bottlenecks associated with dynamic JSON Schema validation in Web APIs, this paper proposes a compilation-based validation paradigm that transforms dynamic interpretation into static compilation of native validators. We introduce the first JSON Schema compilation framework supporting preprocessing of cross-keyword complex constraints, ensuring strict specification compliance while correcting logical flaws present in mainstream tools. Leveraging abstract syntax tree analysis, pattern specialization, finite-state machine generation, and cache-aware code generation, our approach enables end-to-end compilation from schema to highly efficient validators. Experimental evaluation demonstrates an average 10× speedup in validation latency, with up to 100× acceleration in specific scenarios; compilation time remains within seconds to minutes, and runtime incurs zero interpretation overhead.
Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
This study addresses the challenges of maintaining consistency across heterogeneous schema languages—such as JSON Schema, XSD, and SHACL—during multilingual data model evolution, where fragmented converters, variable quality, and information loss impede reliable interoperability. The work proposes a novel approach that models schema languages and black-box converters as nodes and directed edges in a graph, enabling composable and evaluable conversion path orchestration. By integrating graph-based search, quality-aware ranking (combining agent-assisted and human evaluation), and failure backtracking, the method supports automated, reproducible cross-language schema transformation. The resulting open-source toolchain, Schema Conversion Orchestrator, integrated into the MetaConfigurator platform, successfully produced valid outputs for 43 out of 60 real-world tasks and precisely identified missing ecosystem components in the remaining 17, thereby delineating the current boundaries of schema conversion capabilities.
This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.
This work addresses the lack of semantic interoperability in structured data (e.g., JSON, YAML, CSV) within scientific workflows, which hinders consistent interpretation. The authors introduce an RDF authoring view in the MetaConfigurator editor that leverages AI-assisted generation of RML mappings to automatically transform structured data into RDF. The system supports triple editing, SPARQL querying, and knowledge graph visualization. Key innovations include the first integration of large language models for natural language-to-SPARQL translation, bidirectional synchronization between JSON-LD and RDF triples, and ontology-aware IRI auto-completion. Demonstrated on MOF synthesis experiments, the approach successfully converts JSON protocols into semantic knowledge graphs, enabling interactive exploration of relationships between experimental conditions and outcomes, thereby significantly lowering the barrier to adopting Semantic Web technologies.
This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.