🤖 AI Summary
This study addresses the severe Service Level Objective (SLO) violations that arise when local large language model (LLM) servers concurrently serve multiple LoRA adapters, a problem primarily caused by substantial computational overhead and rigid scheduling. To overcome these limitations, this work proposes HALO, a novel scheduling framework that achieves spatial multiplexing of base model and LoRA computations through GPU streaming multiprocessor (SM) partitioning. Furthermore, HALO incorporates an SLO-aware decoupled scheduling mechanism based on per-request slack to effectively mitigate resource contention. Experimental results demonstrate that HALO significantly reduces SLO violation rates while improving overall system throughput, consistently outperforming existing state-of-the-art baselines across all evaluated metrics.
📝 Abstract
As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objectives (SLOs) when serving multiple LoRA-adapted services simultaneously. Existing serving frameworks suffer from severe SLO violations due to the computational overhead of LoRA layers and the rigid nature of batch scheduling. To address this, we propose HALO, a scheduling method tailored for LoRA-assisted on-premise LLM deployment. HALO introduces two key innovations: a spatial multiplexing strategy that overlaps Base and LoRA computations by partitioning GPU Streaming Multiprocessors (SMs), and an SLO-aware scheduler that decouples request execution based on"request-level slack."By prioritizing urgent tasks and utilizing idle budget for traffic shaping, HALO significantly mitigates resource contention. Our evaluation demonstrates that HALO minimizes SLO violations while improving throughput compared to state-of-the-art baselines.