large-scale data processing

Designs and implements scalable systems and pipelines to ingest, clean, transform, aggregate, and analyze very large datasets using distributed and parallel processing techniques. Builds and tunes batch and stream workflows, addressing data partitioning, serialization, resource management, fault tolerance, and performance (throughput, latency, and cost) to ensure reliable, efficient large-scale data processing.

large-scaledataprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling

Dec 15, 2024
JW
Jonathan Will
🏛️ Technische Universitaet Berlin | University of Glasgow

This paper investigates whether autoscaling in Spark serverless environments can improve resource utilization efficiency under fixed hardware constraints—particularly rigid node-level memory-to-CPU ratios. Leveraging fine-grained execution logs from large-scale production Spark batch jobs on Google Dataproc Serverless, we conduct controlled experiments augmented with statistical significance testing and granular resource monitoring. Our analysis, the first at the node level, reveals that current autoscaling mechanisms—constrained by immutable node sizes and static resource allocations—fail to dynamically adapt to workload demands; empirical results show no statistically significant improvement in resource efficiency. The core contribution is the identification and validation of “node-level resource rigidity” as the fundamental bottleneck to resource optimization in serverless Spark. This finding provides critical empirical evidence to guide the design of next-generation elastic schedulers capable of fine-grained, topology-aware resource orchestration.

Auto Resource ScalingBig Data ProcessingResource Optimization

This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.

extreme-scale data processingHigh-Throughput Computingresource utilization

Declarative Data Pipeline for Large Scale ML Services

Aug 20, 2025
YY
Yunzhao Yang
🏛️ Amazon Web Services

To address the joint optimization challenges of performance, maintainability, and collaborative efficiency in large-scale integrated machine learning within distributed data processing systems, this paper proposes Pipes—a declarative, modular data pipeline architecture. Pipes decomposes pipelines into logically encapsulated computation units, implemented atop Apache Spark with standardized interfaces and well-defined component boundaries—departing from conventional microservice paradigms to enable high-performance, maintainable ML pipeline development. In enterprise deployments, Pipes improves development efficiency by 50%, reduces collaborative debugging cycles from weeks to days, achieves 500× scalability, and delivers 10× higher throughput. Academic benchmarks show >5.7× throughput improvement and 99% CPU utilization. Its core contribution is the first deep integration of declarative abstractions with Spark’s native execution model, simultaneously advancing both development methodology and system performance.

Balancing performance with maintainability in large-scale ML servicesIntegrating machine learning capabilities efficiently within Apache SparkReducing communication overhead in collaborative data processing environments

Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges

Oct 29, 2025
GL
Georgios L. Stavrinides
🏛️ Aristotle University of Thessaloniki

Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.

Addressing data locality and parallelism for data-intensive applicationsMeeting QoS requirements like time constraints and energy efficiencyScheduling complex workloads in large-scale distributed systems

To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.

Designing co-operation in hierarchical multi-objective schedulers for stream processingEnhancing load balancing across compute resources for growing application complexityIntegrating new schedulers into existing hierarchies to improve proactive resource management

Latest Papers

What's happening recently
View more

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.

cloud-based data pipelinescost-makespan trade-offsinfrastructure performance

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

iDDS: Intelligent Distributed Dispatch and Scheduling for Workflow Orchestration

Oct 03, 2025
WG
Wen Guan
🏛️ Brookhaven National Laboratory | University of Texas at Arlington | University of Pittsburgh

To address the challenge of efficiently orchestrating and intelligently managing complex, dynamic workflows in large-scale distributed scientific computing, this paper proposes an integrated intelligent workflow system that unifies task scheduling, data movement, and adaptive decision-making. The system supports data-aware execution, conditional logic, and programmable directed acyclic graphs (DAGs), operating in both template-driven and “function-as-a-task” modes. It adopts a modular, message-driven architecture and deeply integrates mainstream middleware—including PanDA and Rucio—while incorporating distributed hyperparameter optimization and AI-assisted modeling. Its cross-experiment, cross-platform design significantly enhances scalability and reproducibility. Deployed in major scientific projects—including ATLAS, the Rubin Observatory, and the Electron-Ion Collider—the system enables high-throughput execution of heterogeneous tasks and reduces operational overhead by over 30%.

Integrating data-aware execution with conditional logic automationOrchestrating large-scale distributed scientific computing workflowsUnifying workload scheduling and data movement across infrastructures

Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications

Nov 15, 2025
MS
Mills Staylor
🏛️ University of Virginia | Biocomplexity Institute and Initiative

To address the high communication overhead and poor scalability of serverless architectures in machine learning–intensive, data-heavy workloads, this paper proposes a high-performance computing (HPC)-inspired serverless framework. Methodologically, it introduces a NAT-traversal direct communication mechanism based on TCP hole punching, implements a lightweight serverless communicator, and integrates the Cylon distributed dataframe library with an FMI-inspired heuristic communication scheduling model. This design enables decentralized, low-latency, high-throughput distributed data processing within cloud-native environments—without centralized coordination. Experimental results demonstrate that the framework achieves over 99% end-to-end performance improvement compared to conventional serverless approaches. Its strong scaling efficiency closely matches that of EC2 instances and dedicated HPC clusters. Notably, this work is the first to achieve near-HPC communication efficiency and scalability in a serverless setting.

Addressing slow communication in serverless functions for large datasetsBridging performance gap between serverless and HPC for data processingImproving distributed data frame performance using direct communication techniques

Hot Scholars

MT

Mike Thelwall

School of Information, Journalism and Communication, The University of Sheffield
scientometricsaltmetricssentiment analysissocial media
AT

Amaury Trujillo

Researcher, IIT-CNR
Social ComputingHuman-Computer InteractionWeb Technologies
GL

Gabriele Lenzini

Interdiscilplinary Centre for Security Reliability and Trust ( SNT) - University of Luxembourg
Sociotechnical Cybersecurity
MK

Max Kleiman-Weiner

University of Washington
Cognitive ScienceCooperationReinforcement LearningConsumer Behavior