Performance and Stability of Barrier Mode Parallel Systems with Heterogeneous and Redundant Jobs

📅 2025-12-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses performance degradation and stability deterioration in barrier-mode parallel systems (e.g., Spark’s Barrier Execution Mode) caused by synchronization barriers. We propose a systematic modeling and analytical framework, introducing the first stochastic process model for $(s,k,l)$-redundant barrier systems and deriving rigorous stability conditions and performance upper bounds using queueing theory. We identify dual-event triggering combined with polling-based scheduling as the dominant source of overhead, and establish analytically tractable performance bounds for mixed workloads—both with and without barrier tasks. The model is validated via distribution fitting, overhead attribution analysis, and empirical measurements on Spark; simulation and measurement results show latency distribution errors under 8%. Our contributions provide a theoretical foundation and quantitative toolkit for designing, optimizing, and ensuring stability in barrier-synchronized parallel systems.

Technology Category

Reasoning under Uncertainty: Stochastic OptimizationPlanning, Routing, and Scheduling: Optimization of Spatio-temporal SystemsConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
In some models of parallel computation, jobs are split into smaller tasks and can be executed completely asynchronously. In other situations the parallel tasks have constraints that require them to synchronize their start and possibly departure times. This is true of many parallelized machine learning workloads, and the popular Apache Spark processing engine has recently added support for Barrier Execution Mode, which allows users to add such barriers to their jobs. These barriers necessarily result in idle periods on some of the workers, which reduces their stability and performance, compared to equivalent workloads with no barriers. In this paper we will consider and analyze the stability and performance penalties resulting from barriers. We include an analysis of the stability of $(s,k,l)$ barrier systems that allow jobs to depart after $l$ out of $k$ of their tasks complete. We also derive and evaluate performance bounds for hybrid barrier systems servicing a mix of jobs, both with and without barriers, and with varying degrees of parallelism. For the purely 1-barrier case we compare the bounds and simulation results to benchmark data from a standalone Spark system. We study the overhead in the real system, and based on its distribution we attribute it to the dual event and polling-driven mechanism used to schedule barrier-mode jobs. We develop a model for this type of overhead and validate it against the real system through simulation.
Problem

Research questions and friction points this paper is trying to address.

Analyzes stability and performance penalties from synchronization barriers in parallel systems.
Models overhead in barrier-mode jobs due to dual event and polling-driven scheduling.
Evaluates performance bounds for hybrid systems with mixed barrier and non-barrier jobs.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Barrier mode analysis for parallel systems stability
Hybrid barrier systems performance bounds derivation
Overhead modeling for dual event polling mechanism
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5