Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of accurately predicting inference overhead in diffusion large model serving, where traditional scalar cost proxies fail to capture the two-dimensional "chunk-step" execution structure. To overcome this limitation, this work pioneers a denoising workload surface modeling mechanism that preserves the two-dimensional structure while weighting heterogeneous costs for precise estimation. Methodologically, it introduces a coarse-to-fine training strategy alongside a chunk-autoregressive generation mechanism, and constructs a lightweight, prompt-only predictor on the CPU to facilitate cross-hardware transferability. Experimental results demonstrate that the proposed approach reduces prediction error by up to 2.5× and decreases end-to-end latency in online chatbot scenarios by up to 1.92×.
📝 Abstract
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.
Problem

Research questions and friction points this paper is trying to address.

Diffusion LLMs
Inference cost prediction
Model serving
Denoising workload
Request scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion LLMs
Denoising Workload Surface
Inference Cost Prediction
Serving Scheduling
Block-autoregressive Generation
🔎 Similar Papers
No similar papers found.