Pie: A Programmable Serving System for Emerging LLM Applications

📅 2025-10-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM serving systems rely on fixed token-generation loops, hindering flexible adaptation to diverse inference strategies and agent-based workflows. Method: We propose Inferlet, the first framework to decouple LLM serving into programmable, fine-grained components—exposing via open APIs KV cache management, generation logic control, and compute-I/O co-scheduling capabilities—thereby delegating generation control to user-defined lightweight programs (inferlets). Inferlets execute securely, efficiently, and in isolation using WebAssembly sandboxing, enabling end-to-end customization without kernel modifications. Results: Experiments show minimal overhead—only 3–12% latency increase on standard generation tasks—while delivering 1.3×–3.4× higher throughput and significantly reduced end-to-end latency on representative agent workflows.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguagePlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Emerging large language model (LLM) applications involve diverse reasoning strategies and agentic workflows, straining the capabilities of existing serving systems built on a monolithic token generation loop. This paper introduces Pie, a programmable LLM serving system designed for flexibility and efficiency. Pie decomposes the traditional generation loop into fine-grained service handlers exposed via an API and delegates control of the generation process to user-provided programs, called inferlets. This enables applications to implement new KV cache strategies, bespoke generation logic, and seamlessly integrate computation and I/O-entirely within the application, without requiring modifications to the serving system. Pie executes inferlets using WebAssembly, benefiting from its lightweight sandboxing. Our evaluation shows Pie matches state-of-the-art performance on standard tasks (3-12% latency overhead) while significantly improving latency and throughput (1.3x-3.4x higher) on agentic workflows by enabling application-specific optimizations.
Problem

Research questions and friction points this paper is trying to address.

Addresses inflexibility in LLM serving systems for diverse reasoning strategies
Decomposes monolithic token generation into programmable fine-grained handlers
Enables custom KV cache strategies and generation logic without system modifications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decomposes generation loop into programmable service handlers
Delegates control to user programs called inferlets
Executes inferlets using WebAssembly for sandboxing
🔎 Similar Papers
No similar papers found.