Score
Designs and implements scalable systems and pipelines to ingest, clean, transform, aggregate, and analyze very large datasets using distributed and parallel processing techniques. Builds and tunes batch and stream workflows, addressing data partitioning, serialization, resource management, fault tolerance, and performance (throughput, latency, and cost) to ensure reliable, efficient large-scale data processing.
This paper investigates whether autoscaling in Spark serverless environments can improve resource utilization efficiency under fixed hardware constraints—particularly rigid node-level memory-to-CPU ratios. Leveraging fine-grained execution logs from large-scale production Spark batch jobs on Google Dataproc Serverless, we conduct controlled experiments augmented with statistical significance testing and granular resource monitoring. Our analysis, the first at the node level, reveals that current autoscaling mechanisms—constrained by immutable node sizes and static resource allocations—fail to dynamically adapt to workload demands; empirical results show no statistically significant improvement in resource efficiency. The core contribution is the identification and validation of “node-level resource rigidity” as the fundamental bottleneck to resource optimization in serverless Spark. This finding provides critical empirical evidence to guide the design of next-generation elastic schedulers capable of fine-grained, topology-aware resource orchestration.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.
To address the joint optimization challenges of performance, maintainability, and collaborative efficiency in large-scale integrated machine learning within distributed data processing systems, this paper proposes Pipes—a declarative, modular data pipeline architecture. Pipes decomposes pipelines into logically encapsulated computation units, implemented atop Apache Spark with standardized interfaces and well-defined component boundaries—departing from conventional microservice paradigms to enable high-performance, maintainable ML pipeline development. In enterprise deployments, Pipes improves development efficiency by 50%, reduces collaborative debugging cycles from weeks to days, achieves 500× scalability, and delivers 10× higher throughput. Academic benchmarks show >5.7× throughput improvement and 99% CPU utilization. Its core contribution is the first deep integration of declarative abstractions with Spark’s native execution model, simultaneously advancing both development methodology and system performance.
Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.
To address insufficient multi-objective load balancing in stream processing systems under complex workloads, this paper proposes a multi-tier collaborative scheduling framework. The framework introduces dynamic inter-layer coordination mechanisms and lightweight interfaces among schedulers, enabling seamless integration of novel scheduling policies. It jointly optimizes computational resource utilization, end-to-end latency, and throughput by integrating multi-objective optimization, distributed resource management, and real-time feedback control. Its key innovation lies in shifting hierarchical scheduling from static decoupling to dynamic collaboration—preserving scalability while significantly enhancing adaptability. Evaluated in Meta’s production environment, the system reliably processes TB-scale data with sub-second latency; it improves critical resource utilization by 27% and reduces tail latency by 41%.
To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.
This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
To address the challenge of efficiently orchestrating and intelligently managing complex, dynamic workflows in large-scale distributed scientific computing, this paper proposes an integrated intelligent workflow system that unifies task scheduling, data movement, and adaptive decision-making. The system supports data-aware execution, conditional logic, and programmable directed acyclic graphs (DAGs), operating in both template-driven and “function-as-a-task” modes. It adopts a modular, message-driven architecture and deeply integrates mainstream middleware—including PanDA and Rucio—while incorporating distributed hyperparameter optimization and AI-assisted modeling. Its cross-experiment, cross-platform design significantly enhances scalability and reproducibility. Deployed in major scientific projects—including ATLAS, the Rubin Observatory, and the Electron-Ion Collider—the system enables high-throughput execution of heterogeneous tasks and reduces operational overhead by over 30%.
To address the high communication overhead and poor scalability of serverless architectures in machine learning–intensive, data-heavy workloads, this paper proposes a high-performance computing (HPC)-inspired serverless framework. Methodologically, it introduces a NAT-traversal direct communication mechanism based on TCP hole punching, implements a lightweight serverless communicator, and integrates the Cylon distributed dataframe library with an FMI-inspired heuristic communication scheduling model. This design enables decentralized, low-latency, high-throughput distributed data processing within cloud-native environments—without centralized coordination. Experimental results demonstrate that the framework achieves over 99% end-to-end performance improvement compared to conventional serverless approaches. Its strong scaling efficiency closely matches that of EC2 instances and dedicated HPC clusters. Notably, this work is the first to achieve near-HPC communication efficiency and scalability in a serverless setting.