Score
Designs and evaluates schemes that divide data, computation, and system resources into partitions—e.g., by key range, graph cut, block, memory region, power domain, or hardware/software boundary—and implements partitioning and compaction methods, indexing layouts, and workload placement algorithms. Work involves building algorithms and tools to assign data and workloads across nodes or domains (including edge/cloud or mixed‑criticality boundaries), and analyzing trade‑offs in locality, load balance, consistency, resource usage, and scalability while specifying the partitioning policies and metadata needed to operate them.
To address the high overhead of dynamic data repartitioning in multi-core HPC systems under time-varying workloads, this paper proposes a lightweight, hierarchical partitioning method jointly driven by geometric and statistical principles. The method integrates space-filling curve ordering, greedy knapsack-based load balancing, and hierarchical data decomposition to support efficient dynamic partitioning of 2D/3D structured grids, point sets, and general graphs. It introduces, for the first time, an adaptive repartitioning mechanism guided by real-time feedback on data distribution, substantially reducing computational and communication overhead in frequently updated scenarios. Implemented via a hybrid parallel programming model (MPI + OpenMP) on modern many-core architectures, experimental results demonstrate a 3.2–5.7× speedup in partitioning time and a load imbalance ratio below 3.1%. This approach provides timely, low-overhead data partitioning support for parallel algorithms in large-scale scientific computing.
This study addresses the practical disparities and co-evolution between high-performance computing (HPC) and edge computing architectures within the cloud continuum. It presents the first large-scale empirical analysis based on 396 real-world, production-grade AWS architectures. Methodologically, we propose a multidimensional, data-driven framework encompassing service topology identification, storage type classification, architectural complexity quantification, and ML service integration statistics. Results reveal systematic differences—and complementary patterns—between HPC and edge architectures across four dimensions: core service composition (e.g., EC2 versus Greengrass/Lambda), storage design paradigms (parallel file systems versus distributed lightweight caches), complexity distributions, and ML embedding strategies. This work delivers the first industry-scale architectural benchmark for the cloud continuum, providing empirically grounded insights and methodological foundations for cross-domain architecture design, resource optimization, and cloud-native convergence of HPC and edge computing.
This study investigates the role of task replication in graph partitioning and DAG scheduling, aiming to substantially reduce or even eliminate inter-processor communication with minimal computational overhead. It presents the first systematic analysis of how replication affects the computational complexity of these two problems and introduces an optimal replication model based on integer linear programming (ILP) alongside an efficient heuristic algorithm tailored for large-scale DAGs. Experimental results demonstrate that, in hypergraph partitioning, communication cost is reduced by 17%–65% on average, with complete elimination in certain scenarios; in DAG scheduling, reductions range from 11.61% to 23.13% on average, reaching as high as 58.17%. These findings confirm the effectiveness and practical value of replication strategies.
Temporal interference arising from shared cache and memory bandwidth among multiple virtual machines (VMs) in mixed-criticality embedded systems undermines timing predictability and system certification. Method: We propose the first configurable, analyzable co-interference mitigation framework. It systematically models and quantifies the coupled effects of cache coloring and memory bandwidth reservation (MBR); employs hardware performance counter (HPC)-driven feedback for configuration optimization; and unifies analysis across multiple shared resources—including IOMMU and interrupt controllers. Contribution/Results: Evaluated on a real multi-core platform, the framework significantly improves worst-case execution time (WCET) predictability and throughput stability. Interference suppression effectiveness increases by 32%, while configuration efficiency improves by 5.8×. The open-source implementation supports industrial-grade deployment in mixed-criticality systems.
This work addresses the joint optimization of communication and computation overheads in distributed computing, where a master node coordinates \(N\) workers to compute a set of subfunctions dependent on \(d\) input files. The problem is modeled as a \(d\)-uniform hypergraph edge partitioning task, and a deterministic Interweaved-Cliques (IC) assignment scheme is proposed. This scheme achieves order-optimal communication load (number of files received per worker) and computation load (number of subfunctions processed per worker) without prior knowledge of the subfunction structure. Leveraging an information-theoretically inspired interwoven clique construction and a deterministic allocation strategy, the method applies to any multi-function decomposition satisfying mild density conditions, requires no file reassignment, and attains order-optimal communication cost \(\Theta(n/N^{1/d})\) and computation cost across a broad range of parameters, yielding a partitioning gain of \(N^{1/d}\).
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This work proposes a highly scalable parallel recursive spectral bisection (parRSB) method tailored for Exascale computing to address the challenge of efficient, high-quality graph partitioning of large-scale spectral element meshes. The approach computes the Fiedler vector in parallel on the mesh dual graph by integrating the Lanczos algorithm with conjugate gradient-based inverse iteration, augmented with several numerical optimization strategies to enhance computational efficiency. Experimental results on the Summit and Frontier supercomputing platforms demonstrate that parRSB significantly reduces communication overhead while preserving partition quality, exhibiting excellent strong and weak scalability as well as superior partitioning speedup.
This work addresses the challenge of simultaneously minimizing worst-case communication overhead and computational load for general functions admitting d-ary decompositions in distributed computing. The authors propose a deterministic Interweaved Clique (IC) assignment framework grounded in combinatorial design theory. This approach circumvents the restrictive existence conditions of Steiner systems, thereby revealing for the first time the fundamental scaling laws of the problem over a significantly broader range of parameters, while permitting modest heterogeneity in workers’ storage loads. The constructed IC scheme achieves communication cost within a constant factor of 4e from the information-theoretic lower bound and maintains order-wise optimal computation load.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.