DPIFrame: A Dual-Level Parallelism Acceleration Framework for CTR Model Inference

📅 2026-06-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of click-through rate (CTR) model inference on GPUs, which stems from a mismatch between their inherently sequential computation patterns and the massively parallel architecture of modern hardware. To overcome this limitation, the authors propose the first dual-level parallel acceleration framework specifically designed for CTR models. The framework innovatively integrates intra- and inter-module parallelism, an efficient multi-table embedding lookup algorithm, and a workload-aware breadth-first stream scheduling mechanism to enable fine-grained GPU resource management. Experimental results demonstrate that the proposed approach achieves up to a 5.83× speedup in inference throughput and reduces embedding latency by up to 23× compared to state-of-the-art frameworks including PyTorch, TorchRec, HugeCTR, and OneFlow.
📝 Abstract
Deep learning technology has enhanced the ability of Click-through rate (CTR) prediction models to learn features and improve prediction accuracy. However, it is challenging to deploy CTR models on GPU smoothly and perform inference efficiently, because there is a huge mismatch between the serial computational pattern and the parallel model structure. In this paper, we propose DPIFrame, the first dual parallelizable framework to accelerate CTR model inference. In DPIFrame, a) a dual parallelizable architecture is proposed to perform parallel CTR model inference in both intra-module and inter-module; b) an efficient multi-table lookup algorithm is presented for embedding operations through anticipating the whole workload in advance; c) a breadth-first stream scheduling strategy is designed for fine-grained management of parallel computation on GPU to further supporting the dual parallel execution. Extensive experiments are conducted on two real-world datasets, and the results highlight that DPIFrame can reduce the embedding latency efficiently by \textbf{23.0$\times$} compared to PyTorch. Compared with PyTorch, TorchRec, HugeCTR, and OneFlow, DPIFrame can achieve state-of-the-art inference performance on GPU with speedups of \textbf{5.83$\times$}, \textbf{4.29$\times$}, \textbf{2.15$\times$}, and \textbf{2.0$\times$}, respectively.
Problem

Research questions and friction points this paper is trying to address.

CTR model inference
GPU acceleration
parallelism
embedding latency
deep learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual-level parallelism
CTR model inference
multi-table lookup
stream scheduling
GPU acceleration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dezhi Yi
College of Computer Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China
Huifeng Guo
Huifeng Guo
Huawei, Harbin Institute of Technology
Recommender SystemDeep LearningData Mining.
K
Kunpeng Xie
College of Computer Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China
Z
Zhaolong Jian
College of Computer Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China
H
Haochi Yu
College of Computer Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China
W
Wenxuan He
College of Computer Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China
Zhenhua Dong
Zhenhua Dong
Noah's ark lab, Huawei Technologies Co., Ltd.
Recommender systemcausal inferencecountrfactual learningtrustworthy AImachine learning
R
Ruiming Tang
Kuaishou Technology Co., Ltd., Beijing 100085, China
Y
Ye Lu
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China; Tianjin Key Laboratory of Network and Data Science Technology, Tianjin 300350, China; Key Laboratory of Data and Intelligent System Security, Ministry of Education, Tianjin 300350, China