big data

Designs and builds systems and workflows for ingesting, storing, processing, and analyzing datasets whose volume, velocity, or variety require distributed or scalable solutions. This includes designing scalable storage architectures, distributed processing and fault-tolerant data-management pipelines, and analytics workflows using big-data technologies.

bigdata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.74
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An Analysis of HPC and Edge Architectures in the Cloud

Aug 02, 2025
SS
Steven Santillan
🏛️ Escuela Superior Politécnica del Litoral | ESPOL

This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.

Analyze HPC and edge architectures in AWS cloud deploymentsAssess architectural complexity and machine learning services usageInvestigate AWS services prevalence and storage systems used

Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges

Oct 29, 2025
GL
Georgios L. Stavrinides
🏛️ Aristotle University of Thessaloniki

Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.

Addressing data locality and parallelism for data-intensive applicationsMeeting QoS requirements like time constraints and energy efficiencyScheduling complex workloads in large-scale distributed systems

Data Management System Analysis for Distributed Computing Workloads

Oct 01, 2025
KH
Kuan-Chieh Hsu
🏛️ Brookhaven National Laboratory | University of Pittsburgh | Carnegie Mellon University | Oak Ridge National Laboratory | University of Massachusetts | SLAC National Accelerator Laboratory

To address systemic inefficiencies—including resource idleness, redundant data transfers, and violations of data locality—arising from misaligned coordination between the PanDA workflow system and the Rucio data management system in the ATLAS experiment, this paper proposes an end-to-end co-optimization framework. We introduce a novel file-level metadata matching algorithm to precisely associate computing tasks with datasets, and integrate log-based tracing, spatiotemporal imbalance analysis, and anomaly pattern detection to construct a fine-grained, holistic view of data access and movement. Our approach is the first to identify, in production, the root causes of cross-system scheduling mismatches, delivering interpretable performance insights. Empirical validation confirms tangible improvements in resource utilization and system resilience, demonstrating the feasibility and effectiveness of the proposed co-design strategies.

Diagnosing systemic inefficiencies in globally distributed data workflowsLinking workflow and data management systems for performance awarenessReducing unnecessary data transfers through coordinated system optimization

This study addresses the challenges AI-driven scientific workflows encounter in high-performance computing (HPC) environments—namely data intensity, resource heterogeneity, frequent iteration cycles, and I/O bottlenecks—which hinder their compatibility with conventional linear pipeline architectures. The work presents the first systematic integration of AI workflow characteristics with HPC system design, proposing a transformative framework tailored for adaptive intelligent computing environments and articulating twelve practical design principles. This framework incorporates key technologies including containerization, job array scheduling, explicit feedback mechanisms, heterogeneous resource management, and small-file I/O optimization, making it particularly well-suited for high-throughput domains such as computational biology. The resulting guidelines offer researchers actionable strategies to substantially enhance the efficiency, portability, and scalability of AI-HPC workflows.

AI-driven workflowsfoundation modelsheterogeneous resource management

WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows

Mar 17, 2025
FL
Fabian Lehmann
🏛️ Humboldt-Universitaet zu Berlin | Technische Universitaet Berlin | Technische Universitaet Darmstadt | University of Glasgow

In dynamic scientific workflows, the decoupling of task scheduling from data movement often assigns tasks to nodes lacking local input data, causing network congestion and execution delays. To address this, we propose the first workflow-aware joint scheduling framework that enables “data-readiness-driven” task placement via predictive pre-staging of intermediate data replicas, supporting dynamic execution plans. Our approach integrates speculative data pre-staging, dependency-aware scheduling, and lightweight storage management, implemented on a Nextflow+Kubernetes prototype. Experiments across 16 synthetic and real-world workflows demonstrate significant reductions in total completion time—up to 94.5% (synthetic) and 53.2% (real), with only bounded, transient storage overhead. The core contribution is the first holistic co-optimization of scheduling and data movement for dynamic workflows, breaking the traditional separation between these concerns.

Improves resource allocation and scalability in cluster environmentsOptimizes task scheduling and data movement in scientific workflowsReduces network congestion and overall workflow runtime

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing distributed data pipeline systems, which require users to explicitly define complete workflow graphs, by proposing a unified planning and scheduling framework that automatically constructs end-to-end persistent pipelines from implicit goal declarations alone. The approach introduces, for the first time, a numeric-domain-independent planner into the context of persistent scheduling, integrating workflow and resource graph modeling, numeric planning, and network interface scheduling to achieve full automation. Experimental results demonstrate the feasibility and scalability of the method: under a single-machine constraint of one hour of CPU time and 30 GB of memory, the system successfully scheduled a linear pipeline spanning eight sites and comprising fourteen components.

automated planningdata pipelinesdistributed workflows

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.

extreme-scale data processingHigh-Throughput Computingresource utilization

This work proposes a serverless real-time data processing framework that integrates the MapReduce programming model with Function-as-a-Service (FaaS) to address the demand for efficient handling of massive, real-time data streams in modern logistics systems. Built on Kubernetes and Knative, the framework employs an event-driven architecture composed of five loosely coupled services for data ingestion, aggregation, and analysis. It leverages Apache Kafka for event transport, Redis for metadata management, and AWS S3 for durable storage. By innovatively combining MapReduce’s batch-processing semantics with FaaS’s elastic scaling capabilities, the system supports on-demand autoscaling and scale-to-zero functionality. Experimental results demonstrate that the proposed approach achieves low-latency processing while significantly improving resource efficiency, thereby fulfilling the requirements of highly elastic and highly available data pipelines.

data aggregationevent-driven processinglogistics systems

This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.

cloud-based data pipelinescost-makespan trade-offsinfrastructure performance