Adaptive Inference Batching using Policy Gradients

📅 2026-07-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing static batching strategies in handling bursty and heterogeneous inference requests, which struggle to dynamically balance throughput and latency. The authors formulate inference batching and routing as a Markov decision process and train reinforcement learning agents—using REINFORCE and PPO algorithms—on a discrete-event simulator grounded in queueing theory and real-world production traces. The agents make dynamic scheduling decisions based on queue states, request types, and GPU availability. Experimental results demonstrate that, in heterogeneous multi-GPU settings, the proposed approach improves throughput by up to 60% and reduces latency by 25% compared to heuristic policies such as round-robin and shortest queue. Under SLA compliance constraints, performance gains reach as high as 348%. Moreover, the agent autonomously learns an effective workload isolation policy that mitigates head-of-line blocking, highlighting the advantages of reinforcement learning in orchestrating complex, multi-resource scheduling scenarios.
📝 Abstract
Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic. We investigate whether reinforcement learning (RL) can learn adaptive batching and routing policies that outperform these heuristics, training REINFORCE and PPO agents on a discrete-event simulator validated against queuing theory and production traces (Azure Functions, BurstGPT). We formulate the problem as an MDP over queue state, request type and GPU availability, evaluating across standard Poisson traffic, extreme bursts, real-world traces and heterogeneous multi-GPU routing. Our central finding is a clear boundary condition for RL's value in systems problems. In single-GPU settings, a well-tuned static batching policy is already near-optimal under Poisson-like arrivals and RL offers only marginal gains (+0.1% to +1.0%). In multi-GPU heterogeneous routing, however, where fast and slow requests compete for shared resources, the agent discovers a workload-segregation policy that eliminates Head-of-Line blocking, yielding a 3.5x (348%) improvement over Round-Robin and a 48% improvement over the strongest heuristic baseline (Shortest-Queue), with 60% higher throughput and 25% lower latency while respecting SLA constraints. The policy generalizes to unseen bursty and real-world traffic despite training only on synthetic Poisson arrivals and an attention-augmented policy network converges roughly 20% faster than an MLP baseline. These results suggest RL's advantage over engineered heuristics concentrates in combinatorial, multi-resource decisions rather than single-resource temporal scheduling, a practical distinction for deciding where learned policies justify their engineering cost in production inference infrastructure.
Problem

Research questions and friction points this paper is trying to address.

inference serving
adaptive batching
heterogeneous workloads
latency-throughput tradeoff
dynamic resource allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Adaptive Batching
Multi-GPU Routing
Head-of-Line Blocking
Inference Serving
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruslan Sharifullin
Department of Computer Science, Stanford University