Score
Designs and implements processes that transform raw information into the structured forms required by downstream systems or models, including crafting prompt and context formats, serializing rules and interaction traces, preserving temporal ordering of histories, and attaching provenance and metadata. Builds templates, serializers, and validation pipelines to ensure input correctness, consistency, and reproducibility.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.
This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.
Proprietary workflow languages (e.g., Smart Forms/Smart Flow) cause vendor lock-in, poor interoperability, and lack of knowledge traceability. To address these issues, this paper proposes an ontology-based, semantic-aware model transformation approach. It employs an RML-driven JSON→RDF/OWL semantic lifting pipeline, integrated with domain ontology alignment, logical reasoning, and declarative mapping rules to enable automated, verifiable M2M transformation from proprietary formats to BPMN 2.0. The key contribution lies in externalizing transformation knowledge as reusable ontologies and rules—supporting explicit control-flow representation, source-code-level traceability, and cross-vendor adaptability. Evaluated on 69 real-world workflows, the method generated 92 BPMN diagrams with a 94.2% success rate. User studies confirm significant improvements in process comprehension, diagnostic efficiency, and team collaboration.
Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.
This work addresses the challenge of reliably translating natural language into industrial-grade, deployable SysMLv2 models. The authors propose an iterative generate-check-repair framework that, for the first time, integrates a production-level SysMLv2 conformance checker directly into the generation process as a control mechanism rather than a post-processing step. By combining large language model (LLM) generation with deterministic diagnostic feedback and targeted repair strategies—and terminating only when zero errors remain—the method achieves perfect compliance. Evaluated across 604 test cases derived from 151 prompts and four distinct LLMs, the approach elevates single-pass generation compliance from 51.16% to 100%, enabling robust, direct translation of natural language specifications into engineering-ready SysMLv2 models.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
This study addresses the scarcity of realistic, shareable, and privacy-safe validation data for large language model (LLM) agents in manufacturing environments that align with actual Manufacturing Execution System (MES) structures. To resolve this, the authors propose a “Template-as-Ontology” approach, wherein a single Python configuration module uniformly defines the domain ontology for both manufacturing simulators and AI analytics tools, ensuring strict alignment of their data schemas. Grounded in the ISA-95/IEC 62264 standards, the framework models 66 entity types and implements a five-layer pipeline—spanning simulation, PostgreSQL storage, CDC/Iceberg lakehouse ingestion, star-schema transformation, and parameterized AI tooling—to architecturally eliminate AI hallucination. Experiments across six industry templates demonstrate that all KPIs remain within prescribed bounds and achieve a 0% hallucination rate under constraints (versus 43% without constraints, p < 10⁻¹²), confirming the method’s efficacy and cross-industry reusability.
This work addresses the limitation of existing text-to-process modeling approaches, which predominantly focus on control flow while neglecting resource and collaboration perspectives, thereby struggling to generate complete multi-party models. To overcome this, the authors propose a resource-aware generative pipeline that systematically incorporates the resource dimension into large language model (LLM)-driven process modeling for the first time. The method automatically constructs BPMN 2.0 collaboration diagrams from natural language descriptions, explicitly capturing organizational pools, role-based lanes, and inter-organizational message events, and employs an orthogonal layout algorithm for automated diagram arrangement. Experimental results across ten business processes and nine LLMs demonstrate that the approach accurately extracts resource-related information, maintains high control-flow quality, and incurs only minimal runtime overhead, advancing generative process modeling toward more collaborative and resource-complete representations.
This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.