🤖 AI Summary
Existing LLM serving systems rely on fixed token-generation loops, hindering flexible adaptation to diverse inference strategies and agent-based workflows.
Method: We propose Inferlet, the first framework to decouple LLM serving into programmable, fine-grained components—exposing via open APIs KV cache management, generation logic control, and compute-I/O co-scheduling capabilities—thereby delegating generation control to user-defined lightweight programs (inferlets). Inferlets execute securely, efficiently, and in isolation using WebAssembly sandboxing, enabling end-to-end customization without kernel modifications.
Results: Experiments show minimal overhead—only 3–12% latency increase on standard generation tasks—while delivering 1.3×–3.4× higher throughput and significantly reduced end-to-end latency on representative agent workflows.
📝 Abstract
Emerging large language model (LLM) applications involve diverse reasoning strategies and agentic workflows, straining the capabilities of existing serving systems built on a monolithic token generation loop. This paper introduces Pie, a programmable LLM serving system designed for flexibility and efficiency. Pie decomposes the traditional generation loop into fine-grained service handlers exposed via an API and delegates control of the generation process to user-provided programs, called inferlets. This enables applications to implement new KV cache strategies, bespoke generation logic, and seamlessly integrate computation and I/O-entirely within the application, without requiring modifications to the serving system. Pie executes inferlets using WebAssembly, benefiting from its lightweight sandboxing. Our evaluation shows Pie matches state-of-the-art performance on standard tasks (3-12% latency overhead) while significantly improving latency and throughput (1.3x-3.4x higher) on agentic workflows by enabling application-specific optimizations.