data processing

Designs, implements, and evaluates pipelines and components that ingest, clean, transform, validate, and serialize data so it is usable for storage, analysis, or downstream algorithms. Work includes parsing and format conversion, data cleaning and normalization, feature extraction and aggregation, batching/streaming, validation and lineage tracking, and related performance and reliability engineering.

dataprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Towards Next Generation Data Engineering Pipelines

Jul 18, 2025
KM
Kevin M. Kramer
🏛️ University of Hagen | University of Regensburg

Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.

Achieving self-awareness and self-adaptationEnabling reactivity to data changesImproving data quality in engineering pipelines

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

PRE-Share Data: Assistance Tool for Resource-aware Designing of Data-sharing Pipelines

Mar 17, 2025
SM
Sepideh Masoudi
🏛️ Technische Universität Berlin

In cross-organizational data sharing, existing multi-pipeline transformation design suffers from low efficiency and severe resource waste under dual constraints of governance compliance and recipient-side adaptability. Method: This paper proposes a reuse-aware pipeline design assistance paradigm that integrates flowchart-based modeling, semantic matching of transformation operations, fine-grained resource consumption modeling, and heuristic configuration optimization. It enables automatic identification of reusable transformation components across pipelines, recommends optimal pipeline structures, and quantifies potential resource savings. Contribution/Results: As the first design assistance framework supporting predictive reporting generation, it achieves, on real-world use cases, an average 37% reduction in computational resource consumption and a 52% reduction in design cycle time, while remaining compatible with self-service data platform deployments.

Designing efficient data-sharing pipelines across organizationsEnsuring compliance with governance policies and recipient requirementsReusing transformation processes to optimize resource consumption

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

This work addresses the long-standing challenge of fan-out and fault-line pitfalls in data transformation caused by granularity mismatches, which traditionally rely on costly runtime testing and manual debugging. We propose the first type-theoretic formal framework for data granularity, modeling granularity relationships as mathematical structures within a type system to enable reasoning over arbitrary data types. By integrating compile-time type-level verification with formal proofs in Lean 4 and large language model assistance, our approach provides end-to-end correctness guarantees with zero runtime overhead. Experimental results demonstrate that the method automatically detects granularity-related errors and reduces verification costs by 98–99%, substantially enhancing the reliability of AI-generated data pipelines.

chasm trapsdata granularitydata transformation correctness

Latest Papers

What's happening recently
View more

This work addresses the challenges posed by the massive and unstructured time-series data generated in industrial cyber-physical systems (CPS), where existing preprocessing approaches rely on ad hoc scripts that suffer from poor readability, reusability, and maintainability. To overcome these limitations, the authors propose and implement CPSLint, the first domain-specific language (DSL) tailored for industrial CPS data preprocessing. CPSLint abstracts common data cleaning and validation operations into a concise and expressive syntax, enabling cross-scenario reuse and significantly improving both data preparation efficiency and team collaboration. The DSL has been open-sourced, and experimental results demonstrate that complex preprocessing tasks can be accomplished in just a few lines of code, substantially reducing redundant development efforts. CPSLint thus establishes a scalable and standardized paradigm for industrial time-series data processing.

Cyber-Physical Systemsdata preparationdata sanitisation

This study addresses the lack of systematic understanding regarding the challenges users encounter in developing and maintaining nf-core standardized bioinformatics pipelines. Conducting the first large-scale empirical analysis, we examined 25,173 GitHub issues and pull requests using BERTopic for topic modeling, Cohen’s δ effect size statistics, and data mining techniques, identifying 13 key problem categories spanning core challenges such as tool development, CI configuration, and containerization debugging. Our findings reveal that 89.38% of reported issues were ultimately resolved, with half addressed within three days. Furthermore, the presence of issue labels and code snippets significantly enhanced resolution efficiency, offering empirical evidence to inform strategies for improving the sustainability and collaborative effectiveness of nf-core pipelines.

GitHub issuesnf-corepipeline maintenance

This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.

code generationexecution feedbackmodel composition

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

Hot Scholars

JG

Jinghuai Gao

Xi'an Jiaotong University
seismic exploration
ZW

Zixin Wen

Carnegie Mellon University
Machine Learning Theory
JL

Jianghao Lin

Shanghai Jiao Tong University
Large Language ModelsAI AgentsRecommender Systems