Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high GPU energy consumption in large language model inference, noting that existing dynamic voltage and frequency scaling (DVFS) approaches overlook the differing sensitivity of attention and feed-forward network (FFN) layers to frequency adjustments across inference stages. To tackle this, the authors propose AFlex, a novel framework that enables independent DVFS control at the attention and FFN layer levels for the first time. AFlex integrates a decoupled architecture with global scheduling and local DVFS management, further enhanced by interleaved pipelining and dynamic micro-batching to minimize pipeline bubbles and communication overhead. Evaluated on Qwen3-32B and Mixtral-8×7B, AFlex reduces per-token energy consumption by up to 49% and 48%, respectively, compared to the state-of-the-art systems, while meeting service-level objectives for time-to-first-token (TTFT) and time-per-output-token (TPOT).
📝 Abstract
Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.
Problem

Research questions and friction points this paper is trying to address.

energy-efficient LLM serving
disaggregated attention–FFN
flexible frequency scaling
operator-level frequency sensitivity
GPU energy consumption
Innovation

Methods, ideas, or system contributions that make the work stand out.

Disaggregated Attention–FFN
Operator-level DVFS
Energy-efficient LLM serving
Flexible frequency scaling
Interleaved A/F pipeline
🔎 Similar Papers
No similar papers found.
C
Cunchen Hu
China Telecom Cloud Computing Research Institute
L
Liangliang Xu
Xidian University
T
Tian Liu
University of Science and Technology of China
M
Min Lyu
University of Science and Technology of China
Yongkun Li
Yongkun Li
University of Science and Technology of China
Storage SystemMemory and File SystemKey-value SystemGraph System
Sa Wang
Sa Wang
Associate Professor, Institute of Computing Technology, CAS
Cloud ComputingOperating Systems
S
Shuo Quan
China Telecom Cloud Computing Research Institute
Y
Yanan Yang
China Telecom Cloud Computing Research Institute
Wenda Tang
Wenda Tang
China Telecom Cloud Computing Research Institute
CloudEnergyMemory optimization
Yiduo Wang
Yiduo Wang
Postdoc, ACFR, University of Sydney
RoboticsPerception
F
Fu Yu
China Telecom Cloud Computing Research Institute
J
Jie Wu
China Telecom Cloud Computing Research Institute