Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.
📝 Abstract
Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers from limited parallelism, irregular computation, and severe load imbalance, preventing efficient execution on GPU supercomputers. We present SparkleDock, a scalable GSO-based docking framework enabling near-real-time flexible docking. We redesign GSO to expose massive fine-grained parallelism at the glowworm-agent level, and restructure the dominant energy scoring computation into a Tensor Core-compatible formulation, enabling efficient execution of irregular pairwise interactions through structured matrix operations. We further introduce a performance-model-driven scheduling for load balancing and out-of-core scaling across GPUs. SparkleDock achieves 9.7 $\times$ and 18.9 $\times$ speedups over LightDock on single A100 and H100 GPU, and delivers over two orders of magnitude acceleration at scale. On 512 GPUs, it reduces docking time from hours to seconds, enabling large-scale, high-fidelity virtual screening previously impractical with flexible docking.
Problem

Research questions and friction points this paper is trying to address.

macromolecular docking
flexible docking
GPU supercomputers
load imbalance
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU acceleration
macromolecular docking
Glowworm Swarm Optimization
Tensor Core
load balancing
Xiangyu Meng
Xiangyu Meng
中国石油大学(华东)
HPC AI
Peng Chen
Peng Chen
RIKEN Center for Computational Science (R-CCS)
HPCGPGPUMachine LearningImage Processing
Mingzhen Li
Mingzhen Li
Institute of Computing Technology, Chinese Academy of Sciences
HPCAI System
J
Jianmin Wang
Department of Computer Science and Engineering, The Chinese University of Hong Kong, China
S
Sen Wang
College of Computer Science and Technology, Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software, China University of Petroleum (East China), Qingdao, China
G
Guangming Tan
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
W
Weile Jia
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
Mohamed Wahib
Mohamed Wahib
RIKEN Center for Computational Science (R-CCS)
High Performance ComputingParallel/Distributed ComputingHigh-Performance AI ...
Tao Luo
Tao Luo
Senior Scientist, Institute of High Performance Computing;Nanyang Technological University
X
Xun Wang
College of Computer Science and Technology, Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software, China University of Petroleum (East China), Qingdao, China