Score
Designs, builds, and analyzes deployment plans and runtime allocation strategies that map software artifacts and their compute, memory, storage, and network requirements onto target hardware, including packing and versioning auxiliary datasets and assigning workloads to devices. Evaluates and optimizes these resource mappings for constraints and objectives such as latency, throughput, accuracy, and utilization, and implements monitoring and adaptation to manage runtime resource use.
This paper addresses the task scheduling challenge in geo-distributed computing, arising from network heterogeneity, heterogeneous resource pricing, and imbalanced computational capacity. It systematically surveys scheduling techniques across four paradigms: cloud, edge, cloud–edge collaboration, and high-performance computing (HPC). The study introduces the first unified taxonomy covering all four environments, grounded in three core objectives—performance, fairness, and fault tolerance—and identifies cross-cutting challenges including cross-domain latency-sensitive scheduling, multi-regional cost optimization, and elastic fault tolerance. Through bibliometric analysis and qualitative comparative evaluation, it classifies and assesses state-of-the-art approaches—including multi-objective optimization, game-theoretic models, reinforcement learning, and heuristic algorithms. The work traces the technical evolution of scheduling research and proposes six future directions: AI-native schedulers, carbon-aware scheduling, among others—thereby providing theoretical foundations and practical guidance for building adaptive, sustainable distributed scheduling systems.
This paper addresses the lack of systematic optimization for CPU and memory resource allocation during the Release phase of cloud-native DevOps. We propose the first pre-deployment offline performance optimization framework for microservices—distinct from mainstream auto-scaling research focused on the Ops phase. Our approach performs fine-grained resource configuration tuning *before* deployment, thereby mitigating auto-scaling failures caused by suboptimal memory provisioning. Methodologically, we integrate Bayesian optimization, statistical experimental design, and a goal-directed factor screening strategy to balance sampling cost and approximation accuracy. Extensive evaluation on the TeaStore benchmark demonstrates that our pre-deployment optimization significantly improves memory suitability and API-level resource utilization. Moreover, it empirically validates the necessity and context-dependent applicability of factor screening under varying optimization objectives.
To address the challenge of dynamic multi-VM scheduling under hardware resource constraints in software-defined vehicles, this paper proposes a workload-aware hypervisor scenario configuration auto-generation framework. Methodologically, it introduces the first integration of domain-knowledge-guided parameter modeling with deep learning to construct a dynamic QoS prediction model; further, it designs an optimization algorithm that synthesizes chip vendors’ BSPs, heuristic rules, and system-level constraints to generate customized resource allocation schemes. The key contributions are: (1) automated and adaptive VM-level resource allocation, and (2) significant improvements in resource utilization and integration efficiency of in-vehicle virtualization systems. Experimental evaluation on real automotive platforms demonstrates a 32% reduction in development cycle time and a 27% average increase in resource utilization.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This work addresses the challenge of efficiently exploring the vast and physically constrained design space of cross-layer heterogeneous systems to support mixed AI and high-performance computing (HPC) workloads. To this end, the authors propose CHASE, a novel framework that decouples hardware architecture design from task mapping. CHASE leverages hierarchical type graphs for system modeling, a topology-aware mapper, and a telemetry-guided optimizer to enable application-driven architecture search under deployment constraints. Experimental results demonstrate that CHASE achieves geometric mean speedups of 6.20× and 2.12× on sparse computing and large language model workloads, respectively, while reducing mapping time by 60.5% on average and converging to near-global-optimal solutions within 64 iterations.
Traditional operating systems suffer from poor scalability on many-core processors and low parallel efficiency due to their inability to perceive application semantics. To address this, we propose NetworkedOS—a novel application-aware, networked OS architecture. Our approach leverages compile-time dynamic instruction dependency analysis to construct a multi-layer network model that explicitly captures runtime dependencies among applications, the kernel, and hardware. We further design an overlapping graph partitioning algorithm to jointly optimize parallel execution and inter-core communication overhead, and implement a runtime process affinity mapping scheduler. Crucially, NetworkedOS breaks the conventional “black-box” OS assumption regarding application semantics for the first time. Experimental evaluation shows that NetworkedOS achieves a 7.11× speedup over Linux on a 128-core system and a 2.01× improvement over Barrelfish on a 64-core system, significantly enhancing scalability and resource utilization under large-scale parallel workloads.
Existing performance analysis tools struggle to simultaneously capture temporal dynamics and a holistic view of performance bottlenecks: Roofline models neglect time evolution, while profilers and tracers obscure theoretical performance limits. This work proposes campaign diagrams—a novel visualization framework that uniquely integrates temporal phases with multidimensional resource utilization, including computational throughput, memory bandwidth, data traffic, and latency. Campaign diagrams can be generated from analytical models, simulations, or profiling data, concurrently displaying both theoretical performance ceilings and achieved performance. The approach uncovers cross-phase optimization opportunities often missed by conventional tools, such as counterintuitive cases where enhancing low-intensity operators improves end-to-end performance. Validated on low-rank GEMM and Mamba workloads, the method successfully identifies potential for operator fusion and pipeline optimizations, demonstrating its efficacy in diagnosing deep-rooted performance bottlenecks.
This work addresses the persistent challenge researchers face in establishing reproducible, GPU-ready computational environments—even after securing cloud or on-premise GPU resources. To lower this barrier, the authors propose and implement the first lightweight “adaptation layer” tailored for scientific computing, built upon k3s and Coder and integrated with a GitHub-driven CI/CD pipeline. This system enables fully automated deployment of interactive research environments within five minutes. The study further introduces a quantitative evaluation framework that defines key metrics such as deployment latency and reproducibility. Empirical validation in real-world research scenarios demonstrates that the approach substantially reduces both environment setup time and user friction, thereby enhancing research efficiency and reproducibility.
This study addresses the lack of systematic optimization in cloud data pipelines with respect to cost, execution time, and resource utilization, particularly in multi-tenant and industrial settings where research remains limited. Through a comprehensive systematic literature review, the work establishes a unified classification framework for optimization objectives that encompasses both single- and multi-cloud environments as well as batch and stream processing paradigms. The analysis synthesizes existing approaches and identifies critical research gaps, including insufficient support for multi-tenancy, inadequate multi-cloud coordination, and a scarcity of real-world deployment validation. By clarifying the core objectives and technical pathways for optimizing cloud data pipelines, this paper provides a theoretical foundation and clear direction for future research in this domain.
This work addresses the persistent challenge of inconsistent development and execution environments faced by researchers operating across heterogeneous computing platforms—ranging from laptops and workstations to supercomputers and cloud infrastructures. To overcome this, the authors propose a modular and portable software ecosystem featuring a unified command-line interface that enables seamless orchestration and execution of scientific workflows. The system ensures cross-platform consistency, reproducibility, and scalability, thereby streamlining computational research across diverse hardware configurations. Its practical efficacy has been demonstrated through successful integration into the plan4res project under the European Union’s Horizon 2020 initiative, where it effectively supported complex, large-scale scientific workflows in varied computing environments.
This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.