Cross-Facility LLM Pre-training on HPC: Elastic Aggregation, Data Leasing, and Queue-Aware Placement

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of fragmented academic computing resources and cross-facility large model pre-training by proposing a distributed training framework that integrates multiple global supercomputing centers. Methodologically, building upon the DiLoCo dual-loop architecture, it introduces elastic Nesterov outer-step optimization, a DARL heartbeat-based data leasing protocol, and a privilege-free queue-aware placement mechanism to enable resilient aggregation and efficient scheduling of intercontinental resources. Experimental results demonstrate that the system achieves fault-tolerant pre-training with zero data loss while reducing overhead to 3.1%, significantly shortening training cycles. This work provides an efficient and viable new paradigm for decentralized, large-scale LLM training.
πŸ“ Abstract
Academic compute is fragmented: allocations are granted per facility, and facilities differ in accelerators and software stacks, schedule jobs independently, and share neither a network nor a filesystem. We present a system that pools such allocations to pre-train a single language model across three supercomputers on two continents, up to 7,400km apart: Snellius (NVIDIA H100), LUMI and Frontier (both AMD MI250X). It combines (i) DiLoCo-style two-loop training with an elastic, token-weighted Nesterov outer step for which zero, one or many live sites are all normal states; (ii) DARL, a data-leasing protocol whose heartbeat-backed leases guarantee that, within an epoch, no sample is trained twice or lost under crashes, late joins and work stealing; and (iii) queue-aware placement driven by unprivileged sbatch --test-only probes. Training Qwen3-0.6B on C4 for 20,000 optimizer steps, the three-site run reaches a held-out perplexity of 34.7, against 28.2 for a centralised baseline. Per-round overhead (weight exchange and checkpointing) stays near 110 s regardless of the number of local steps H, so its share of wall-clock time falls from 32% at H=100 to 5.7% at H=1,000 and 3.1% at H=2,000. In a 23.8 h three-site run with four site departures, 1.2% of granted data blocks were reclaimed and none was duplicated or lost. In an idealized queue-model projection, queue-aware placement shortens time-to-target by 18-43% compared with waiting for all sites to be allocated. Cross-site pre-training is thus operational rather than competitive: it turns fragmented allocations into one training run at a measured cost.
Problem

Research questions and friction points this paper is trying to address.

Cross-facility pre-training
Fragmented compute allocations
Heterogeneous accelerators
Distributed LLM training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-facility pre-training
Elastic aggregation
Data leasing
Queue-aware placement
Distributed optimization
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Z
Zarè Palanciyan
SURF, Amsterdam, The Netherlands
T
Thomas van Osch
SURF, Amsterdam, The Netherlands
D
Douwe van der Wal
SURF, Amsterdam, The Netherlands
O
Olivera Kotevska
Oak Ridge National Laboratory, Oak Ridge, TN, USA
T
Tim Kok
SURF, Amsterdam, The Netherlands