🤖 AI Summary
This study addresses the challenge of intra-rack resource fragmentation in multi-tenant TPU clusters, where small-to-medium jobs frequently leave idle capacity difficult to utilize. To overcome this, we propose a defragmentation scheduler based on Optical Circuit Switching (OCS) that employs a locality-first strategy combined with on-demand cross-rack splicing. Furthermore, this work introduces a novel low-overhead logical ring construction method that transcends the limitations of conventional coarse-grained, full-rack scheduling by converting fragmented capacity into schedulable resources. Evaluations via TPU v8-scale SuperPod simulations demonstrate that the proposed approach substantially enhances cluster scheduling capability, achieving superior performance compared to both local placement and full-rack baselines.
📝 Abstract
Large-scale AI training clusters increasingly use optical circuit switching (OCS) to reconfigure rack-level interconnects and create elastic accelerator slices. In multi-tenant TPU-style clusters, however, small and medium jobs often leave partial free capacity stranded inside racks. Although the aggregate free capacity may be sufficient for a new job, it cannot be used by local placement or coarse full-rack stitching. This paper presents RingStitch, an OCS-based defragmentation scheduler that turns fragmented rack capacity into schedulable resources. RingStitch follows a local-first policy, stitches compact cross-rack fragments only when needed, and orders the selected racks into a low-cost logical ring. Simulations on a TPU 8t-like SuperPod model show that RingStitch improves schedulability over local and full-rack baselines.