AI-assisted JSON Schema Creation and Mapping

📅 2025-08-07
📈 Citations: 0
Influential: 0
📄 PDF

career value

200K/year
🤖 AI Summary
Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.

Technology Category

Application Category

📝 Abstract
Model-Driven Engineering (MDE) places models at the core of system and data engineering processes. In the context of research data, these models are typically expressed as schemas that define the structure and semantics of datasets. However, many domains still lack standardized models, and creating them remains a significant barrier, especially for non-experts. We present a hybrid approach that combines large language models (LLMs) with deterministic techniques to enable JSON Schema creation, modification, and schema mapping based on natural language inputs by the user. These capabilities are integrated into the open-source tool MetaConfigurator, which already provides visual model editing, validation, code generation, and form generation from models. For data integration, we generate schema mappings from heterogeneous JSON, CSV, XML, and YAML data using LLMs, while ensuring scalability and reliability through deterministic execution of generated mapping rules. The applicability of our work is demonstrated in an application example in the field of chemistry. By combining natural language interaction with deterministic safeguards, this work significantly lowers the barrier to structured data modeling and data integration for non-experts.
Problem

Research questions and friction points this paper is trying to address.

Lack of standardized models in many domains
Difficulty in JSON Schema creation for non-experts
Challenges in mapping heterogeneous data formats
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-assisted JSON Schema creation via LLMs
Hybrid natural language and deterministic techniques
Scalable schema mapping for heterogeneous data formats
🔎 Similar Papers
2024-05-27arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
F
Felix Neubauer
Institute for Parallel and Distributed Systems, University of Stuttgart, Stuttgart, Germany
B
Benjamin Uekermann
Institute for Parallel and Distributed Systems, University of Stuttgart, Stuttgart, Germany
J
Jürgen Pleiss
Institute of Biochemistry and Technical Biochemistry, University of Stuttgart, Stuttgart, Germany