locality-aware scheduling

Design and implement scheduling and load‑balancing policies, runtime mechanisms, and analyses that assign tasks or place data based on object and cache locality to minimize remote-state accesses and contention. Build locality analyses and cache-locality optimizations that maintain throughput and fairness across execution workers and adapt to dynamic access patterns.

locality-awarescheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions

Jul 11, 2024
HZ
Hailiang Zhao
🏛️ Zhejiang University | Nanyang Technological University

This work addresses data-locality-aware task assignment and online scheduling for distributed job execution under unknown job arrival sequences, aiming to minimize job completion time. We propose the Optimal Balanced Task Assignment (OBTA) algorithm, which theoretically improves the approximation ratio of the water-filling algorithm; design a more efficient Replica Deletion (RD) heuristic that significantly reduces computational overhead while preserving accuracy; and introduce a job reordering mechanism based on Shortest Estimated Time First (SEFT) to enhance throughput and response efficiency. Leveraging combinatorial optimization modeling, online algorithm design, and trace-driven evaluation, our approach achieves substantial reductions in average job completion time over real-world workloads compared to state-of-the-art baselines. The RD heuristic outperforms classical water-filling in both speed and effectiveness, and SEFT-based reordering further boosts overall scheduling performance.

Achieving data-locality-aware task assignment and schedulingMinimizing job completion times without future arrival knowledgeReducing computational overhead while maintaining performance

This work addresses the performance degradation caused by frequent workload migration in modern multicore systems, which disrupts microarchitectural locality—particularly in chiplet-based architectures where cross-LLC-domain execution intensifies interference. To mitigate this, the authors propose a user-space-guided kernel scheduling mechanism that prioritizes spatial locality. The approach dynamically allocates compact CPU affinity sets as soft hints by online estimation of CPU demand and explicit awareness of LLC topology, eschewing rigid partitioning or fully shared policies. This strategy significantly enhances locality while maintaining high resource utilization. Empirical evaluation demonstrates throughput improvements of 12% on chiplet systems and 3% on non-chiplet systems, alongside 3%–7% higher per-gigabyte throughput due to reduced memory footprint.

cache localitychiplet architectureCPU scheduling

ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling

Mar 14, 2025
AF
Alessandro Fogli
🏛️ Imperial College London | Aalto University | TU Munich

To address intra-chiplet memory access imbalance and inefficient task scheduling caused by partitioned L3 caches in chiplet-based CPUs, this paper proposes a lightweight adaptive runtime system that jointly optimizes task scheduling, memory allocation, and performance monitoring. Our approach introduces a chiplet-aware fine-grained task migration mechanism and a hardware-topology-aware memory allocation strategy—overcoming the limitations of conventional NUMA optimizations in chiplet architectures. It integrates chiplet-aware heuristic scheduling, a user-space lightweight concurrency model (supporting suspension/resumption and cross-chiplet task migration), and real-time performance monitoring. Experimental evaluation across diverse memory-intensive parallel applications demonstrates an average 1.7× speedup, a 22% improvement in L3 cache hit rate, and a 35% reduction in cross-chiplet memory access latency.

Address memory contention in chiplet-based CPUsImprove task scheduling across NUMA domainsOptimize cache utilization for parallel applications

The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture

Mar 20, 2025
MY
Mert Yildiz
🏛️ University of Rome Sapienza

This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”

Comparing simple vs. complex dispatching policies for workload scheduling.Exploring the impact of parallelism and system architecture on performance.Optimizing job response time in cloud computing clusters.

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

Latest Papers

What's happening recently
View more

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

This work addresses the challenge that task scheduling under strong scaling is often constrained by task granularity, where scheduling overhead can dominate performance as parallelism increases, yet a systematic understanding of how algorithmic dependency structures affect scheduling efficiency remains lacking. The paper proposes a novel framework that characterizes granularity based on the dependency topology of task graphs, attributing the growth of scheduling overhead to dependency structure rather than problem size—a distinction made for the first time. Building on this insight, the authors develop a predictive model for strong-scaling limits and derive rules for selecting appropriate scheduling strategies. Through task graph analysis and overhead modeling, the approach accurately explains both gradual and abrupt scaling breakdowns observed across diverse parallel workloads, enabling informed automatic selection between static and dynamic scheduling without exhaustive empirical testing.

dependency topologydynamic schedulingscheduling overhead

In dynamic multi-tenant environments, programmable caching engines such as CacheLib often suffer from performance degradation, memory inefficiency, and unfair service allocation due to rigid configuration schemes, insufficient runtime adaptability, and the absence of quality-of-service (QoS) guarantees. This work presents the first systematic empirical evaluation of CacheLib under fluctuating workloads across a range of configurations, uncovering its critical bottlenecks and limitations. The study not only quantifies the shortcomings of current designs in terms of fairness and efficiency but also provides clear guidance for future enhancements aimed at improving dynamic adaptability, QoS support, and programmability in caching systems.

cache adaptabilitydynamic workloadsmulti-tenant

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This study addresses the limitations of static cache replacement and prefetching policies in conventional processors, which struggle to maintain optimal performance across diverse execution phases. For the first time, it systematically evaluates the potential of dynamic policy selection by analyzing 490 execution phases from 49 benchmark programs using the ChampSim simulator. The results demonstrate that static policies incur an average IPC loss of 1.54%, whereas dynamically switching between two carefully selected policies reduces this loss to just 0.11%. Moreover, such dynamic adaptation achieves near-ideal performance in 52.65% of the phases, closely approaching the theoretical upper bound. This work validates the efficacy of dynamic policy switching and establishes a new paradigm for enhancing single-threaded performance.

cache replacementdynamic policy selectionout-of-order pipeline

Hot Scholars

IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
YZ

Yanfeng Zhang

Northeastern University, China
Database SystemsMachine Learning Systems
YY

Yuichi Yoshida

National Institute of Informatics
Theoretical Computer Science