high concurrency

Designs and implements software architectures, components, and operational practices that enable systems to correctly and efficiently handle very large numbers of simultaneous requests, connections, or tasks. Work includes defining concurrency models and synchronization, choosing event-driven or thread-based execution, implementing non‑blocking/lock‑free algorithms and resource management, and applying load balancing, scaling, and performance testing to meet throughput and latency targets under heavy concurrent load.

highconcurrency

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Minimize Your Critical Path with Combine-and-Exchange Locks

Nov 12, 2025
SK
Simon König
🏛️ University of Stuttgart

Existing user-space coroutine/fiber synchronization mechanisms implicitly assume kernel scheduling, introducing unnecessary latency on critical paths and limiting high-concurrency throughput. This paper proposes Combine-and-Exchange Scheduling (CES), a novel synchronization paradigm for purely user-space cooperative scheduling. CES eliminates cross-thread overhead by retaining critical sections on the same thread during lock contention, while dynamically redistributing parallelizable tasks to idle threads. Crucially, it co-designs user-space synchronization primitives with the scheduler to fully bypass kernel intervention. Experimental evaluation demonstrates that CES achieves up to 3× higher throughput on application-level benchmarks and up to 8× speedup on microbenchmarks—significantly outperforming state-of-the-art user-space synchronization approaches.

Improving throughput for coroutine-based parallel applicationsOptimizing scheduling for contended critical sections across threadsReducing critical path delays in userspace synchronization primitives

On the Duality of Task and Actor Programming Models

Aug 22, 2025
RY
Rohan Yadav
🏛️ Stanford University | NVIDIA

Task-based and actor-based programming models face a fundamental trade-off between developer productivity and runtime performance, hindering their joint adoption in distributed heterogeneous systems. Method: We establish a formal duality between the two models and propose a unified modeling framework with low-overhead scheduling and communication optimizations. Our approach integrates explicit and implicit parallelism within the Realm/Legion task runtime, enabling fine-grained dependency-aware scheduling, zero-copy inter-node communication, and lightweight task migration. Contribution/Results: Experiments show that Realm reduces runtime overhead by 1.7–5.3× and improves strong scaling by 1.3–5.0×, achieving end-to-end performance competitive with mature actor systems (e.g., Charm++ and MPI). This work is the first to rigorously formalize and empirically validate the duality of task and actor models—both theoretically and in practice—thereby establishing a foundation for high-productivity, high-performance programming paradigms in distributed heterogeneous environments.

Bridging performance gap between task and actor systemsExploring duality between task and actor programming modelsReducing overheads in task-based runtimes for better performance

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.

extreme-scale data processingHigh-Throughput Computingresource utilization

Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs

Oct 23, 2024
FK
Ferdinand Kossmann
🏛️ MIT | Databricks

This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.

GPU OptimizationLarge Language ModelsResource Allocation

Latest Papers

What's happening recently
View more

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the challenge of effectively evaluating the trade-offs between data consistency and coordination overhead among distributed transaction patterns—such as Saga and TCC—in business logic-intensive microservice systems prior to production deployment. The authors propose a lightweight microservice simulator grounded in Domain-Driven Design (DDD), which, for the first time, integrates DDD aggregate root modeling with multiple transaction models to decouple business logic from communication and transactional infrastructure. The framework supports configurable deployment topologies and network constraints, enabling seamless transitions from centralized to fully distributed architectures while providing a deterministic verification environment. Empirical evaluation on complex multi-aggregate systems quantifies the performance, coordination overhead, and resilience of different transaction models, substantially reducing development costs and facilitating left-shifted architectural validation.

architectural simulationconsistency modelsdistributed transactions