🤖 AI Summary
This study addresses the underutilization of GPU memory in large language model (LLM) serving, where conventional coarse-grained memory reclamation disrupts fine-tuning and fails to satisfy inference latency service-level objectives (SLOs). To this end, this work proposes a fine-grained memory sharing system that enables inference to instantaneously preempt activation memory from fine-tuning tasks while preserving training continuity through recomputation during backpropagation. By integrating parameter-efficient fine-tuning with dynamic memory management, the system achieves sub-second reclamation of individual activations and ensures safe preemption under asynchronous CPU-GPU execution and tensor parallelism. Experimental results demonstrate that the proposed approach improves fine-tuning throughput by 1.9× to 3.3× while maintaining an inference SLO attainment rate exceeding 99.7%.
📝 Abstract
Large language model (LLM) serving scales its replica count with the request load, yet GPU memory still stands idle inside the replicas. Adding a replica takes minutes, while the memory that a replica needs changes within seconds. Even instant autoscaling could not return this idle memory, because the smallest unit that it can remove is a whole replica. Colocating parameter-efficient fine-tuning (PEFT) with inference can use this memory, but inference must be able to reclaim it within seconds, before requests that wait for memory exceed their latency service-level objective (SLO). Existing colocation systems either keep the tuning memory resident or let inference reclaim it at the coarse granularity of a whole training sample. Each such reclamation also discards the running tuning step. To address these limitations, we present MOLT, a fine-grained memory sharing system that lets inference reclaim the memory of individual activations that a running tuning step has saved for its backward pass. The step continues, and its backward pass recomputes those activations. Inference reclaims only memory that no in-flight GPU work can still access, even under CPU--GPU asynchrony and tensor parallelism. On four model deployments (24B--70B) across H100 SXM and B200 GPUs under trace-driven workloads, MOLT keeps inference SLO attainment at or above 99.7% and completes 1.9--3.3x the tuning work of discard-based memory sharing.