data engineering

Designs, builds, and maintains systems and pipelines that collect, ingest, transform, validate, store, and serve data at scale; implements ETL/ELT, batch and streaming processing, data modeling, metadata and schema management, and data quality checks. Automates deployment, monitoring, and operational tooling for data storage and compute (databases, warehouses, lakes, message systems, processing frameworks) to ensure reliable, performant, and governed data delivery to downstream consumers.

dataengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.86
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

This work proposes a hierarchical meta-agent architecture to overcome the static and rigid nature of traditional data processing pipelines, which struggle to autonomously monitor and optimize themselves post-deployment. The framework integrates three core components—planning, agent orchestration, and a monitoring feedback loop—to enable dynamic construction, execution, and iterative refinement of end-to-end workflows. It introduces context-aware optimization, adaptive workload partitioning, and progressive sampling mechanisms, facilitating agent reuse and seamless integration with external tools. These innovations significantly enhance system flexibility and scalability. Experimental results demonstrate the framework’s effectiveness and practicality in automatically constructing, continuously monitoring, and adaptively optimizing diverse data processing tasks.

adaptive pipeline optimizationautonomous data processingcontext-aware data processing

Towards Next Generation Data Engineering Pipelines

Jul 18, 2025
KM
Kevin M. Kramer
🏛️ University of Hagen | University of Regensburg

Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.

Achieving self-awareness and self-adaptationEnabling reactivity to data changesImproving data quality in engineering pipelines

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

Automated Planning for Optimal Data Pipeline Instantiation

Mar 16, 2025
LR
Leonardo Rosa Amado
🏛️ Pontifical Catholic University of Rio Grande do Sul | University of Aberdeen | Johannes Kepler University Linz | Sapienza University of Rome | SAP Labs

This work addresses the problem of efficient cluster deployment for data pipelines in datacenters. We propose a novel method that jointly optimizes computational resource allocation and operator scheduling. The problem is formulated as a planning problem with action costs, encoded in PDDL, to minimize end-to-end execution time while explicitly modeling data transfer overhead, operator execution requirements, and distributed resource constraints. Our key contribution is the introduction of a heuristic planning strategy guided by dataflow graph connectivity—marking the first systematic application of automated planning techniques to pipeline instantiation optimization. Experimental evaluation demonstrates significant improvements in compute-communication co-scheduling efficiency: across multiple benchmark scenarios, our approach reduces average end-to-end execution time by 23.7% compared to state-of-the-art baselines.

Minimize communication and execution overheadOptimize data pipeline deployment in clustersPropose heuristics to reduce total execution time

Latest Papers

What's happening recently
View more

Enterprise data warehouse task delivery involves complex workflows that existing large models and agents struggle to support in production settings due to insufficient dependency awareness, lifecycle management, and platform evolution capabilities. This work proposes the first end-to-end automated delivery agent framework, which coordinates hierarchical agents to orchestrate warehouse-specific skills, validates artifacts before and after execution through lifecycle-aware controls, and employs a trajectory-driven skill evolution mechanism for continuous optimization. Key innovations include a dependency-aware orchestration scheme, artifact governance within closed-loop execution, and a real-trajectory-based skill refinement approach. Deployed at scale on Tencent Cloud WeData, the system serves 3,600 monthly active users, supports 18,240 delivery sessions per month, achieves an end-to-end success rate of 87.2%, and enables 73.5% autonomous submissions. A/B testing demonstrates a reduction in median delivery time from 228 to 23 minutes and engineer effort from 95 to 11 minutes.

artifact lifecycle controldata warehouse deliverydependency-aware orchestration

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

This work addresses the critical challenge that while autonomous agents enhance operational efficiency, their failures can lead to sudden and irreversible consequences. To mitigate this risk, the paper introduces, for the first time, the concept of an “agent data environment,” which reimagines traditional passive data systems by constructing an active execution substrate that integrates files, APIs, applications, and system states. This architecture simultaneously enables capability enhancement and enforces safety constraints through embedded mechanisms for proactive intervention and assurance. By doing so, it not only amplifies agent effectiveness but also effectively bounds the impact of potential failures, thereby establishing a foundational framework for highly reliable autonomous automation.

agentic automationautonomous agentsdata environments

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

Hot Scholars

SN

Subigya Nepal

Assistant Professor, Computer Science, University of Virginia
Human-Centered AIMental HealthUbiquitous ComputingDigital Health
CF

Chaoyou Fu

Nanjing University
Multimodal LLMLLMBiometrics
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
QJ

Qing Jiang

PhD student, South China University of Technology
Computer VisionOpen-set Object Detection