data pipeline engineering

Designs, builds, and analyzes end-to-end automated data pipelines that generate, ingest, transform, validate, and serve datasets across batch, streaming/real-time, distributed, training, inference, and simulation workflows. Implements pipeline integration, orchestration, scripting, monitoring, and optimization to meet performance and latency targets, preserve data semantics during preprocessing, and support interactive and automated experimental or domain-specific analysis pipelines.

datapipelineengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

Towards Next Generation Data Engineering Pipelines

Jul 18, 2025
KM
Kevin M. Kramer
🏛️ University of Hagen | University of Regensburg

Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.

Achieving self-awareness and self-adaptationEnabling reactivity to data changesImproving data quality in engineering pipelines

Automated Planning for Optimal Data Pipeline Instantiation

Mar 16, 2025
LR
Leonardo Rosa Amado
🏛️ Pontifical Catholic University of Rio Grande do Sul | University of Aberdeen | Johannes Kepler University Linz | Sapienza University of Rome | SAP Labs

This work addresses the problem of efficient cluster deployment for data pipelines in datacenters. We propose a novel method that jointly optimizes computational resource allocation and operator scheduling. The problem is formulated as a planning problem with action costs, encoded in PDDL, to minimize end-to-end execution time while explicitly modeling data transfer overhead, operator execution requirements, and distributed resource constraints. Our key contribution is the introduction of a heuristic planning strategy guided by dataflow graph connectivity—marking the first systematic application of automated planning techniques to pipeline instantiation optimization. Experimental evaluation demonstrates significant improvements in compute-communication co-scheduling efficiency: across multiple benchmark scenarios, our approach reduces average end-to-end execution time by 23.7% compared to state-of-the-art baselines.

Minimize communication and execution overheadOptimize data pipeline deployment in clustersPropose heuristics to reduce total execution time

To address the challenges of parallel scheduling, opaque execution states, poor result reproducibility, and inadequate auditability when managing hundreds to thousands of Snakemake/Nextflow pipelines in large-scale bioinformatics analyses, this paper proposes a lightweight command-line orchestration framework. Built in Python and integrated with SQLite or PostgreSQL, it enables unified pipeline launching across heterogeneous workflows, real-time status monitoring, fine-grained log collection, automated result ingestion into databases, and comprehensive lifecycle metric logging—including runtime, resource consumption, and failure points. It introduces a novel CLI paradigm that supports cross-pipeline collaborative monitoring and reproducibility assurance without modifying existing workflow code. Experimental evaluation demonstrates a 42% improvement in multi-project throughput, significantly enhancing observability, auditability, and reproducibility in large-scale bioinformatics analysis.

Automate provisioning and evaluation of bioinformatics pipelinesCoordinate bulk processing of multiple datasets efficientlyMonitor and record pipeline metrics for reproducibility

Declarative Data Pipeline for Large Scale ML Services

Aug 20, 2025
YY
Yunzhao Yang
🏛️ Amazon Web Services

To address the joint optimization challenges of performance, maintainability, and collaborative efficiency in large-scale integrated machine learning within distributed data processing systems, this paper proposes Pipes—a declarative, modular data pipeline architecture. Pipes decomposes pipelines into logically encapsulated computation units, implemented atop Apache Spark with standardized interfaces and well-defined component boundaries—departing from conventional microservice paradigms to enable high-performance, maintainable ML pipeline development. In enterprise deployments, Pipes improves development efficiency by 50%, reduces collaborative debugging cycles from weeks to days, achieves 500× scalability, and delivers 10× higher throughput. Academic benchmarks show >5.7× throughput improvement and 99% CPU utilization. Its core contribution is the first deep integration of declarative abstractions with Spark’s native execution model, simultaneously advancing both development methodology and system performance.

Balancing performance with maintainability in large-scale ML servicesIntegrating machine learning capabilities efficiently within Apache SparkReducing communication overhead in collaborative data processing environments

Latest Papers

What's happening recently
View more

Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.

action tracedata analysisend-to-end workflow

This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.

debugginglarge language modelsreproducibility

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the limitation of existing distributed data pipeline systems, which require users to explicitly define complete workflow graphs, by proposing a unified planning and scheduling framework that automatically constructs end-to-end persistent pipelines from implicit goal declarations alone. The approach introduces, for the first time, a numeric-domain-independent planner into the context of persistent scheduling, integrating workflow and resource graph modeling, numeric planning, and network interface scheduling to achieve full automation. Experimental results demonstrate the feasibility and scalability of the method: under a single-machine constraint of one hour of CPU time and 30 GB of memory, the system successfully scheduled a linear pipeline spanning eight sites and comprising fourteen components.

automated planningdata pipelinesdistributed workflows

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
ZX

Zhiheng Xi

Fudan University
LLM ReasoningLLM-based Agents
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc