🤖 AI Summary
This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.
📝 Abstract
Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers from limited parallelism, irregular computation, and severe load imbalance, preventing efficient execution on GPU supercomputers. We present SparkleDock, a scalable GSO-based docking framework enabling near-real-time flexible docking. We redesign GSO to expose massive fine-grained parallelism at the glowworm-agent level, and restructure the dominant energy scoring computation into a Tensor Core-compatible formulation, enabling efficient execution of irregular pairwise interactions through structured matrix operations. We further introduce a performance-model-driven scheduling for load balancing and out-of-core scaling across GPUs. SparkleDock achieves 9.7 $\times$ and 18.9 $\times$ speedups over LightDock on single A100 and H100 GPU, and delivers over two orders of magnitude acceleration at scale. On 512 GPUs, it reduces docking time from hours to seconds, enabling large-scale, high-fidelity virtual screening previously impractical with flexible docking.