data platforms

Designs, builds, and operates scalable systems that collect, ingest, store, process, and serve data, including pipelines, storage and compute layers, metadata/catalogs, APIs, and monitoring/deployment components. Implements data governance, security, reliability, performance tuning, and integration with analytics and downstream consumers.

dataplatforms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

An Analysis of HPC and Edge Architectures in the Cloud

Aug 02, 2025
SS
Steven Santillan
🏛️ Escuela Superior Politécnica del Litoral | ESPOL

This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.

Analyze HPC and edge architectures in AWS cloud deploymentsAssess architectural complexity and machine learning services usageInvestigate AWS services prevalence and storage systems used

This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.

cloud-based data pipelinescost-makespan trade-offsinfrastructure performance

Existing data systems lack embedded accountability mechanisms when supporting decision-making, yielding outputs that, while efficient and accurate, often fail to guarantee justifiability, constraint compliance, and operational feasibility. This work proposes RAIDS, a novel framework that treats accountability not as post-hoc metadata but as an integral execution semantics. RAIDS introduces responsibility contracts as a first-class operator abstraction and enforces end-to-end accountability through responsibility state propagation and preservation mechanisms spanning the entire pipeline from data processing to decision generation. The framework formally defines responsibility-preserving objectives and establishes a full-stack infrastructure for accountable intelligence, encompassing execution, optimization, provenance, and evaluation. Furthermore, it articulates a comprehensive research agenda for responsibility-aware data systems, laying the theoretical foundation for decision systems that are trustworthy, controllable, and auditable.

actionabilityconstraint satisfactiondecision infrastructure

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This study addresses the growing complexity of operating large-scale computing infrastructures, such as high-performance computing (HPC) systems, where existing operational data analytics (ODA) frameworks struggle to effectively support multi-layered, distributed graph processing ecosystems. The work provides a systematic review of core ODA components, benchmarks prevailing approaches, and proposes a novel holistic ODA framework that integrates the distributed graph processing hierarchy introduced by Sherif Sak et al., thereby extending the functional scope of Netti et al.’s prior work. By unifying fine-grained monitoring, ODA architecture, and graph processing system design, the proposed framework significantly enhances structural integrity and functional extensibility, markedly improving operational efficiency. Furthermore, it illuminates key research directions for ODA in high-performance computing environments.

High-Performance ComputingLarge-scale Computing InfrastructuresOperational Data Analytics

Latest Papers

What's happening recently
View more

To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.

Designing co-operation in hierarchical multi-objective schedulers for stream processingEnhancing load balancing across compute resources for growing application complexityIntegrating new schedulers into existing hierarchies to improve proactive resource management

This work proposes a hierarchical meta-agent architecture to overcome the static and rigid nature of traditional data processing pipelines, which struggle to autonomously monitor and optimize themselves post-deployment. The framework integrates three core components—planning, agent orchestration, and a monitoring feedback loop—to enable dynamic construction, execution, and iterative refinement of end-to-end workflows. It introduces context-aware optimization, adaptive workload partitioning, and progressive sampling mechanisms, facilitating agent reuse and seamless integration with external tools. These innovations significantly enhance system flexibility and scalability. Experimental results demonstrate the framework’s effectiveness and practicality in automatically constructing, continuously monitoring, and adaptively optimizing diverse data processing tasks.

adaptive pipeline optimizationautonomous data processingcontext-aware data processing

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

This work addresses the limitation of existing distributed data pipeline systems, which require users to explicitly define complete workflow graphs, by proposing a unified planning and scheduling framework that automatically constructs end-to-end persistent pipelines from implicit goal declarations alone. The approach introduces, for the first time, a numeric-domain-independent planner into the context of persistent scheduling, integrating workflow and resource graph modeling, numeric planning, and network interface scheduling to achieve full automation. Experimental results demonstrate the feasibility and scalability of the method: under a single-machine constraint of one hour of CPU time and 30 GB of memory, the system successfully scheduled a linear pipeline spanning eight sites and comprising fourteen components.

automated planningdata pipelinesdistributed workflows