format input data

Designs and implements processes that transform raw information into the structured forms required by downstream systems or models, including crafting prompt and context formats, serializing rules and interaction traces, preserving temporal ordering of histories, and attaching provenance and metadata. Builds templates, serializers, and validation pipelines to ensure input correctness, consistency, and reproducibility.

formatinputdata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.

format compliancelarge language modelssoftware engineering

Ontology-Driven Model-to-Model Transformation of Workflow Specifications

Nov 17, 2025
FA
Francisco Abreu
🏛️ Instituto Superior Técnico, Universidade de Lisboa | INESC-ID

Proprietary workflow languages (e.g., Smart Forms/Smart Flow) cause vendor lock-in, poor interoperability, and lack of knowledge traceability. To address these issues, this paper proposes an ontology-based, semantic-aware model transformation approach. It employs an RML-driven JSON→RDF/OWL semantic lifting pipeline, integrated with domain ontology alignment, logical reasoning, and declarative mapping rules to enable automated, verifiable M2M transformation from proprietary formats to BPMN 2.0. The key contribution lies in externalizing transformation knowledge as reusable ontologies and rules—supporting explicit control-flow representation, source-code-level traceability, and cross-vendor adaptability. Evaluated on 69 real-world workflows, the method generated 92 BPMN diagrams with a 94.2% success rate. User studies confirm significant improvements in process comprehension, diagnostic efficiency, and team collaboration.

Addressing vendor lock-in by enabling interoperability between workflow systemsPreserving semantic traceability during model transformation through ontology alignmentTransforming proprietary workflow specifications to standard BPMN format

AI-assisted JSON Schema Creation and Mapping

Aug 07, 2025
FN
Felix Neubauer
🏛️ University of Stuttgart

Current structured data modeling and cross-format schema mapping lack accessible, low-threshold tools—particularly hindering non-expert users. This paper proposes a hybrid approach synergizing large language models (LLMs) with deterministic rule-based processing: LLMs interpret natural-language requirements to generate or refine JSON Schema, while a verifiable rule engine performs high-precision, scalable schema mapping across multiple formats (JSON, CSV, XML, YAML). The method is implemented in the open-source tool MetaConfigurator, supporting visual schema modeling and automated code generation. Empirical evaluation in the chemistry domain demonstrates substantial reductions in modeling barriers, significant improvements in schema construction efficiency and mapping accuracy, and—critically—the first end-to-end data schema engineering solution that is natural-language-driven, flexible, and formally reliable.

Challenges in mapping heterogeneous data formatsDifficulty in JSON Schema creation for non-expertsLack of standardized models in many domains

Latest Papers

What's happening recently
View more

This work addresses the challenge of reliably translating natural language into industrial-grade, deployable SysMLv2 models. The authors propose an iterative generate-check-repair framework that, for the first time, integrates a production-level SysMLv2 conformance checker directly into the generation process as a control mechanism rather than a post-processing step. By combining large language model (LLM) generation with deterministic diagnostic feedback and targeted repair strategies—and terminating only when zero errors remain—the method achieves perfect compliance. Evaluated across 604 test cases derived from 151 prompts and four distinct LLMs, the approach elevates single-pass generation compliance from 51.16% to 100%, enabling robust, direct translation of natural language specifications into engineering-ready SysMLv2 models.

Conformance CheckingExecutable ModelsModel-Based Systems Engineering

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

This study addresses the scarcity of realistic, shareable, and privacy-safe validation data for large language model (LLM) agents in manufacturing environments that align with actual Manufacturing Execution System (MES) structures. To resolve this, the authors propose a “Template-as-Ontology” approach, wherein a single Python configuration module uniformly defines the domain ontology for both manufacturing simulators and AI analytics tools, ensuring strict alignment of their data schemas. Grounded in the ISA-95/IEC 62264 standards, the framework models 66 entity types and implements a five-layer pipeline—spanning simulation, PostgreSQL storage, CDC/Iceberg lakehouse ingestion, star-schema transformation, and parameterized AI tooling—to architecturally eliminate AI hallucination. Experiments across six industry templates demonstrate that all KPIs remain within prescribed bounds and achieve a 0% hallucination rate under constraints (versus 43% without constraints, p < 10⁻¹²), confirming the method’s efficacy and cross-industry reusability.

data privacydomain schemamanufacturing AI validation

This work addresses the limitation of existing text-to-process modeling approaches, which predominantly focus on control flow while neglecting resource and collaboration perspectives, thereby struggling to generate complete multi-party models. To overcome this, the authors propose a resource-aware generative pipeline that systematically incorporates the resource dimension into large language model (LLM)-driven process modeling for the first time. The method automatically constructs BPMN 2.0 collaboration diagrams from natural language descriptions, explicitly capturing organizational pools, role-based lanes, and inter-organizational message events, and employs an orthogonal layout algorithm for automated diagram arrangement. Experimental results across ten business processes and nine LLMs demonstrate that the approach accurately extracts resource-related information, maintains high control-flow quality, and incurs only minimal runtime overhead, advancing generative process modeling toward more collaborative and resource-complete representations.

BPMN collaboration diagramcontrol-flowmulti-collaborative process

This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.

abstract notationconceptual modelinglow-level constructs

Hot Scholars

XM

Xiaofeng Mao

Alibaba Group
Computer VisionAdversarial Machine Learning
KZ

Kaipeng Zhang

Shanghai AI Laboratory
LLMMultimodal LLMsAIGC
YH

Yao Hu

浙江大学
Machine Learning
LK

Lecheng Kong

PhD Student, Washington University in St.Louis
graph learningmachine learning