EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses resource imbalance and downstream request starvation caused by the encoding phase in multimodal large language model serving. We propose an encoding-aware decoupled architecture that repositions encoding as a system control point, synergistically optimizing the encoding, prefill, and decoding pipelines through runtime scheduling and hybrid auto-configuration. The framework incorporates load-adaptive micro-batching, partial offloading, and dynamic streaming multiprocessor (SM) partitioning, further enhanced by Tree-structured Parzen Estimator (TPE) Bayesian optimization and fine-grained GPU sharing techniques to transcend the limitations of conventional text-based LLM serving. Compared with NVIDIA Dynamo and vLLM under identical service-level objectives (SLOs), our approach achieves up to 4.3× higher throughput while significantly improving the balance of GPU utilization.
📝 Abstract
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Disaggregated Serving
Encode-Prefill-Decode Pipeline
Resource Allocation
GPU Utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal LLM Serving
Disaggregated Architecture
Micro-batching
SM Partitioning
Bayesian Optimization
🔎 Similar Papers
No similar papers found.