🤖 AI Summary
This study addresses the performance bottleneck in distributed quantum circuit simulation caused by inefficient three-dimensional torus partitioning geometries. By conducting 39-qubit simulations on the Fugaku supercomputer and employing micro-benchmarks to evaluate the effects of partition geometry and rank density on execution time, this work reveals that network latency, rather than communication volume, constitutes the primary bottleneck, while quantifying the performance penalty induced by suboptimal aspect ratios. The results demonstrate that near-cubic partitions achieve a 1.73× to 2.31× speedup over flattened configurations with reduced energy consumption, approximately halving the optimized solution time. These findings provide critical guidance for topology mapping in large-scale quantum simulations.
📝 Abstract
In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73-2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime $\propto a^{0.39}$, $r = 0.72$). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free all-to-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads.