🤖 AI Summary
This study addresses the challenge of accurately predicting inference overhead in diffusion large model serving, where traditional scalar cost proxies fail to capture the two-dimensional "chunk-step" execution structure. To overcome this limitation, this work pioneers a denoising workload surface modeling mechanism that preserves the two-dimensional structure while weighting heterogeneous costs for precise estimation. Methodologically, it introduces a coarse-to-fine training strategy alongside a chunk-autoregressive generation mechanism, and constructs a lightweight, prompt-only predictor on the CPU to facilitate cross-hardware transferability. Experimental results demonstrate that the proposed approach reduces prediction error by up to 2.5× and decreases end-to-end latency in online chatbot scenarios by up to 1.92×.
📝 Abstract
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.