Template-as-Ontology: Configurable Synthetic Data Infrastructure for Cross-Domain Manufacturing AI Validation

📅 2026-05-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of realistic, shareable, and privacy-safe validation data for large language model (LLM) agents in manufacturing environments that align with actual Manufacturing Execution System (MES) structures. To resolve this, the authors propose a “Template-as-Ontology” approach, wherein a single Python configuration module uniformly defines the domain ontology for both manufacturing simulators and AI analytics tools, ensuring strict alignment of their data schemas. Grounded in the ISA-95/IEC 62264 standards, the framework models 66 entity types and implements a five-layer pipeline—spanning simulation, PostgreSQL storage, CDC/Iceberg lakehouse ingestion, star-schema transformation, and parameterized AI tooling—to architecturally eliminate AI hallucination. Experiments across six industry templates demonstrate that all KPIs remain within prescribed bounds and achieve a 0% hallucination rate under constraints (versus 43% without constraints, p < 10⁻¹²), confirming the method’s efficacy and cross-industry reusability.
📝 Abstract
LLarge language model (LLM)-based AI agents deployed in manufacturing environments require populated, schema-correct data for validation, yet production MES data is proprietary, privacy-encumbered, and vendor-specific. This paper introduces the Template-as-Ontology principle: a single Python configuration module (700-770 lines, 45 validated exports) serves simultaneously as the specification for a time-stepped manufacturing simulator and as the runtime domain schema for AI analytics tools, producing alignment by construction rather than integration. We formally define the domain template as a typed relational configuration schema and prove that structural alignment between simulation and tool layers is guaranteed by single-source consumption. A five-layer pipeline--simulation, PostgreSQL, CDC/Iceberg lakehouse, star schema, and 12 parameterized AI tools--generates causally coherent, MES-shaped data spanning 66 entity types across four operational domains mapped to ISA-95/IEC 62264. We validate the architecture with six industry templates (aerospace, pharma, automotive, electronics, beverages, warehousing) running on identical framework code. Calibration experiments (60 runs, 10 seeds per template) confirm parametric controllability: observed KPIs fall within configured ranges across all templates. A controlled hallucination experiment (72 tool invocations, Qwen3-32B) demonstrates that ontology-constrained parameters eliminate tool-parameter fabrication (0% constrained vs. 43% unconstrained hallucination rate for the evaluated model, Fisher's exact test p < 10^-12); the 0% constrained rate is an architectural guarantee that holds for any model. The framework provides a reusable data layer for discrete manufacturing AI validation.
Problem

Research questions and friction points this paper is trying to address.

synthetic data
manufacturing AI validation
domain schema
data privacy
MES data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Template-as-Ontology
synthetic data generation
manufacturing AI validation
ontology-constrained LLMs
schema alignment by construction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Grama Chethan
Siemens Digital Industries Software, Plano, TX 75024, United States