data infrastructure

Designs, builds, and operates the systems and components that ingest, store, transform, catalog, serve, secure, and monitor data at scale—including pipelines, databases, warehouses/lakes, streaming platforms, metadata/catalog services, and data access APIs. Develops data models, schemas, ETL/ELT and orchestration workflows, and analyzes the stack’s performance, reliability, cost, security, and governance to meet operational and compliance requirements.

datainfrastructure

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Data Management System Analysis for Distributed Computing Workloads

Oct 01, 2025
KH
Kuan-Chieh Hsu
🏛️ Brookhaven National Laboratory | University of Pittsburgh | Carnegie Mellon University | Oak Ridge National Laboratory | University of Massachusetts | SLAC National Accelerator Laboratory

To address systemic inefficiencies—including resource idleness, redundant data transfers, and violations of data locality—arising from misaligned coordination between the PanDA workflow system and the Rucio data management system in the ATLAS experiment, this paper proposes an end-to-end co-optimization framework. We introduce a novel file-level metadata matching algorithm to precisely associate computing tasks with datasets, and integrate log-based tracing, spatiotemporal imbalance analysis, and anomaly pattern detection to construct a fine-grained, holistic view of data access and movement. Our approach is the first to identify, in production, the root causes of cross-system scheduling mismatches, delivering interpretable performance insights. Empirical validation confirms tangible improvements in resource utilization and system resilience, demonstrating the feasibility and effectiveness of the proposed co-design strategies.

Diagnosing systemic inefficiencies in globally distributed data workflowsLinking workflow and data management systems for performance awarenessReducing unnecessary data transfers through coordinated system optimization

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

This work proposes a hierarchical meta-agent architecture to overcome the static and rigid nature of traditional data processing pipelines, which struggle to autonomously monitor and optimize themselves post-deployment. The framework integrates three core components—planning, agent orchestration, and a monitoring feedback loop—to enable dynamic construction, execution, and iterative refinement of end-to-end workflows. It introduces context-aware optimization, adaptive workload partitioning, and progressive sampling mechanisms, facilitating agent reuse and seamless integration with external tools. These innovations significantly enhance system flexibility and scalability. Experimental results demonstrate the framework’s effectiveness and practicality in automatically constructing, continuously monitoring, and adaptively optimizing diverse data processing tasks.

adaptive pipeline optimizationautonomous data processingcontext-aware data processing

Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.

composable data systemsdata contractsmulti-language lakehouse

Can AI autonomously build, operate, and use the entire data stack?

Dec 08, 2025
AA
Arvind Agarwal
🏛️ IBM Research

Current AI systems only assist specific roles in isolated data management tasks, failing to achieve end-to-end autonomous governance across the enterprise data stack lifecycle. Method: This paper proposes a novel paradigm of autonomous data systems based on multi-agent collaboration, integrating intelligent agents, autonomous decision-making, AI-driven governance, and quality optimization—embedded within modern data stack architectures to enable dynamic, closed-loop data architecture design, integration, monitoring, and evolution. Contribution/Results: The study demonstrates the feasibility of fully autonomous data systems, identifies critical challenges—including cross-layer semantic alignment, trustworthy autonomy boundaries, and human-AI co-governance—and establishes a scalable, auditable, human-AI shared, self-sustaining data ecosystem framework. These contributions provide a foundational methodology and technical roadmap for next-generation highly autonomous data infrastructure.

AI autonomously manages entire data lifecycleBuilds self-sufficient systems for humans and AIShifts from component AI to holistic automation

Latest Papers

What's happening recently
View more

Enterprise data warehouse task delivery involves complex workflows that existing large models and agents struggle to support in production settings due to insufficient dependency awareness, lifecycle management, and platform evolution capabilities. This work proposes the first end-to-end automated delivery agent framework, which coordinates hierarchical agents to orchestrate warehouse-specific skills, validates artifacts before and after execution through lifecycle-aware controls, and employs a trajectory-driven skill evolution mechanism for continuous optimization. Key innovations include a dependency-aware orchestration scheme, artifact governance within closed-loop execution, and a real-trajectory-based skill refinement approach. Deployed at scale on Tencent Cloud WeData, the system serves 3,600 monthly active users, supports 18,240 delivery sessions per month, achieves an end-to-end success rate of 87.2%, and enables 73.5% autonomous submissions. A/B testing demonstrates a reduction in median delivery time from 228 to 23 minutes and engineer effort from 95 to 11 minutes.

artifact lifecycle controldata warehouse deliverydependency-aware orchestration

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.

cloud-based data pipelinescost-makespan trade-offsinfrastructure performance

This study addresses the lack of systematic methodologies for selecting data architectures in modern organizations grappling with vast, heterogeneous data environments. To this end, it proposes the DATER conceptual framework, which establishes a unified taxonomy of technical requirements and systematically examines the historical evolution, core characteristics, and applicability boundaries of six prominent data architectures: data warehouses, data lakes, lakehouses, data fabrics, and data meshes. Through conceptual modeling and multidimensional comparative analysis, the framework clarifies overlaps and distinctions among these architectures, articulating their respective strengths and limitations. By offering a structured evaluation tool, DATER significantly enhances the strategic alignment and contextual appropriateness of data architecture design for both researchers and practitioners.

data architecturedata integrationdata management

Hot Scholars

SC

Samuel Cahyawijaya

Cohere
Low-Resource NLPUnderrepresented LanguagesMultilingualCosslingual
SS

Saber Saberian

MSc, Simon Fraser University
Applied Machine LearningDrug DiscoveryBioinformatics
HA

Hossein Azizpour

Associate Professor, KTH (Royal Institute of Technology)
Computer VisionDeep LearningLife SciencesSustainable Development
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery