🤖 AI Summary
This study addresses the limitation of static load balancing in isolating backend servers that exhibit performance degradation—such as persistently returning HTTP 500 errors—without fully failing, which leads to elevated client-side error rates. The work presents the first systematic investigation into the feasibility of deploying open-source large language models (LLMs) within a real-world load-balancing control plane. Leveraging HAProxy and Prometheus telemetry updated every 10 seconds, the system dynamically quarantines faulty nodes via constrained API calls. Evaluations across 15 open-source LLMs spanning dense, mixture-of-experts (MoE), and sparse architectures reveal that models with approximately 3 billion active parameters represent a capability threshold for effective scheduling. Above this threshold, non-inference-based models reduce 5xx errors by 88%, albeit at the cost of 2.6–2.8× higher tail latency; enabling inference, however, increases token consumption tenfold and degrades control responsiveness.
📝 Abstract
Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server returning HTTP 500s until an operator intervenes. We ask whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers through guardrailed calls to the HAProxy Data Plane API. On a reproducible benchmark with a persistent structural fault built into roughly one-third of a heterogeneous fleet, we sweep 15 open-weight models across five families (0.35B to 35B total parameters; dense, mixture-of-experts, and efficient-sparse architectures), reasoning modes, fleet scales of 3 to 9 backends, and two routing algorithms, totaling 240 runs. We find a capability threshold near 3B active parameters. Below it, LLM policies are typically unreliable and sometimes worse than no policy; above it, every model, regardless of architecture, saturates near an 88% reduction in client-perceived 5xx errors over the static baseline. The threshold is approximate: Gemma 4 E2B clears it with 2B active parameters, while the dense 3B Granite 4.0 Micro does not. The availability gain has costs. Draining concentrates load onto surviving servers, inflating tail latency 2.6 to 2.8 times, and enabling reasoning multiplies token spend roughly tenfold, overrunning the control interval and degrading effectiveness. The efficient operating point is a supra-threshold model in its cheapest non-reasoning mode, wrapped inside deterministic guardrails.