data production

Design, build, and evaluate systems and pipelines that produce and manage datasets, including instrumentation for data capture, ETL and preprocessing, annotation and labeling workflows, synthetic-data generation and augmentation, metadata and versioning, storage and access interfaces, and quality-control and privacy/compliance mechanisms. Create reproducible, documented data assets suitable for downstream analysis or model training and measure or analyze their throughput, fidelity, bias, and provenance.

dataproduction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

Towards Next Generation Data Engineering Pipelines

Jul 18, 2025
KM
Kevin M. Kramer
🏛️ University of Hagen | University of Regensburg

Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.

Achieving self-awareness and self-adaptationEnabling reactivity to data changesImproving data quality in engineering pipelines

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

This work addresses the limited effectiveness of synthetic data in data-scarce clinical scenarios—such as intraoperative radiotherapy for breast cancer—where success hinges on the identification and management of critical attributes. The authors propose an attribute-driven synthetic data engineering approach that reframes validity as an engineerable task of attribute lifecycle management. Collaborating closely with oncologists, they systematically develop a synthetic data framework tailored to high-sensitivity medical domains, spanning requirement definition, formal modeling, privacy-preserving validation, and process evolution. Beyond uncovering core challenges in synthetic data engineering for data-scarce software systems, this study advances automated software engineering by introducing mechanisms for collaborative attribute specification and continuous evolution.

data scarcityprivacy constraintsproperty-driven

Latest Papers

What's happening recently
View more

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

This work addresses the challenge of reconciling high throughput and low query latency in traditional ETL pipelines when processing continuously arriving fresh data, where unpredictable preprocessing operations often create bottlenecks. The authors propose Fluid ETL Pipelines, which introduce, for the first time, an elastic and non-blocking preprocessing mechanism that decouples data ingestion from transformation. By dynamically scheduling preprocessing tasks based on resource availability and user interest—without blocking data ingestion—and leveraging preemptible computing resources such as Amazon Spot instances, the approach significantly reduces operational costs. Experimental results demonstrate that Fluid ETL Pipelines substantially improve the efficiency of exploring fresh data, offering a novel direction for accelerating real-time queries and enabling adaptive preprocessing management.

data preprocessing routinesETL pipelinesfresh data exploration

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Scientific data often require extensive manual curation before being usable for scientific AI, lacking a unified framework for automated conversion, readiness assessment, provenance tracking, and agent integration. This work proposes REDI, an open-source framework that automatically transforms raw scientific data into AI-ready formats through a five-stage, fully traceable pipeline—ingestion, preprocessing, transformation, structuring, and output—while exposing the resulting workflows as callable skills for AI agents. REDI is the first framework to unify these capabilities; its companion tool, SetGo, ensures FAIR compliance and enables automatic catalog publishing. Leveraging parallel distributed processing and I/O performance profiling, REDI demonstrates effectiveness across climate science, proteomics, materials science, and nuclear fusion, with the climate use case achieving near-ideal strong scaling up to 100 nodes on the Frontier supercomputer.

automated transformationdata readinessFAIR compliance