Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure

📅 2025-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large-scale CGYRO parameter scans for magnetically confined fusion suffer from high inter-node communication overhead and memory redundancy due to distributed execution across many compute nodes. To address this, we propose XGYRO—the first cooperative execution framework that treats ensemble simulations as a unified computational unit. Its core innovation lies in exploiting structural consistency across collision operator tensors among different simulations, enabling cross-task tensor sharing and global buffer redistribution—thereby transcending conventional per-simulation optimization paradigms. Built upon CGYRO, XGYRO integrates distributed-memory optimization with tensor reuse techniques, substantially reducing per-simulation memory footprint and aggregate communication volume. Experiments on hundred-scale parameter scans demonstrate 40–60% reduction in communication overhead, significantly improving scalability and computational efficiency of ensemble simulations. XGYRO thus provides an efficient infrastructure for exploring high-dimensional fusion physics parameter spaces.

Technology Category

Constraint Satisfaction and Optimization: Distributed CSP/OptimizationSearch and Optimization: Distributed SearchMachine Learning: Ensemble Methods

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Federated Web and WoT systems, including distributed, federated and edge-based data processingSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
First-principles fusion plasma simulations are both compute and memory intensive, and CGYRO is no exception. The use of many HPC nodes to fit the problem in the available memory thus results in significant communication overhead, which is hard to avoid for any single simulation. That said, most fusion studies are composed of ensembles of simulations, so we developed a new tool, named XGYRO, that executes a whole ensemble of CGYRO simulations as a single HPC job. By treating the ensemble as a unit, XGYRO can alter the global buffer distribution logic and apply optimizations that are not feasible on any single simulation, but only on the ensemble as a whole. The main saving comes from the sharing of the collisional constant tensor structure, since its values are typically identical between parameter-sweep simulations. This data structure dominates the memory consumption of CGYRO simulations, so distributing it among the whole ensemble results in drastic memory savings for each simulation, which in turn results in overall lower communication overhead.
Problem

Research questions and friction points this paper is trying to address.

Reducing HPC communication costs in CGYRO ensembles
Sharing collisional tensor structure across simulations
Optimizing memory usage in fusion plasma simulations
Innovation

Methods, ideas, or system contributions that make the work stand out.

XGYRO executes ensemble CGYRO simulations jointly
Shares collisional constant tensor structure globally
Reduces memory usage and communication overhead significantly
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5