Score
Designs and implements software architectures, components, and operational practices that enable systems to correctly and efficiently handle very large numbers of simultaneous requests, connections, or tasks. Work includes defining concurrency models and synchronization, choosing event-driven or thread-based execution, implementing non‑blocking/lock‑free algorithms and resource management, and applying load balancing, scaling, and performance testing to meet throughput and latency targets under heavy concurrent load.
Existing user-space coroutine/fiber synchronization mechanisms implicitly assume kernel scheduling, introducing unnecessary latency on critical paths and limiting high-concurrency throughput. This paper proposes Combine-and-Exchange Scheduling (CES), a novel synchronization paradigm for purely user-space cooperative scheduling. CES eliminates cross-thread overhead by retaining critical sections on the same thread during lock contention, while dynamically redistributing parallelizable tasks to idle threads. Crucially, it co-designs user-space synchronization primitives with the scheduler to fully bypass kernel intervention. Experimental evaluation demonstrates that CES achieves up to 3× higher throughput on application-level benchmarks and up to 8× speedup on microbenchmarks—significantly outperforming state-of-the-art user-space synchronization approaches.
Task-based and actor-based programming models face a fundamental trade-off between developer productivity and runtime performance, hindering their joint adoption in distributed heterogeneous systems. Method: We establish a formal duality between the two models and propose a unified modeling framework with low-overhead scheduling and communication optimizations. Our approach integrates explicit and implicit parallelism within the Realm/Legion task runtime, enabling fine-grained dependency-aware scheduling, zero-copy inter-node communication, and lightweight task migration. Contribution/Results: Experiments show that Realm reduces runtime overhead by 1.7–5.3× and improves strong scaling by 1.3–5.0×, achieving end-to-end performance competitive with mature actor systems (e.g., Charm++ and MPI). This work is the first to rigorously formalize and empirically validate the duality of task and actor models—both theoretically and in practice—thereby establishing a foundation for high-productivity, high-performance programming paradigms in distributed heterogeneous environments.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.
This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This work addresses the challenge of effectively evaluating the trade-offs between data consistency and coordination overhead among distributed transaction patterns—such as Saga and TCC—in business logic-intensive microservice systems prior to production deployment. The authors propose a lightweight microservice simulator grounded in Domain-Driven Design (DDD), which, for the first time, integrates DDD aggregate root modeling with multiple transaction models to decouple business logic from communication and transactional infrastructure. The framework supports configurable deployment topologies and network constraints, enabling seamless transitions from centralized to fully distributed architectures while providing a deterministic verification environment. Empirical evaluation on complex multi-aggregate systems quantifies the performance, coordination overhead, and resilience of different transaction models, substantially reducing development costs and facilitating left-shifted architectural validation.