🤖 AI Summary
This work addresses the inefficiency of click-through rate (CTR) model inference on GPUs, which stems from a mismatch between their inherently sequential computation patterns and the massively parallel architecture of modern hardware. To overcome this limitation, the authors propose the first dual-level parallel acceleration framework specifically designed for CTR models. The framework innovatively integrates intra- and inter-module parallelism, an efficient multi-table embedding lookup algorithm, and a workload-aware breadth-first stream scheduling mechanism to enable fine-grained GPU resource management. Experimental results demonstrate that the proposed approach achieves up to a 5.83× speedup in inference throughput and reduces embedding latency by up to 23× compared to state-of-the-art frameworks including PyTorch, TorchRec, HugeCTR, and OneFlow.
📝 Abstract
Deep learning technology has enhanced the ability of Click-through rate (CTR) prediction models to learn features and improve prediction accuracy. However, it is challenging to deploy CTR models on GPU smoothly and perform inference efficiently, because there is a huge mismatch between the serial computational pattern and the parallel model structure. In this paper, we propose DPIFrame, the first dual parallelizable framework to accelerate CTR model inference. In DPIFrame, a) a dual parallelizable architecture is proposed to perform parallel CTR model inference in both intra-module and inter-module; b) an efficient multi-table lookup algorithm is presented for embedding operations through anticipating the whole workload in advance; c) a breadth-first stream scheduling strategy is designed for fine-grained management of parallel computation on GPU to further supporting the dual parallel execution. Extensive experiments are conducted on two real-world datasets, and the results highlight that DPIFrame can reduce the embedding latency efficiently by \textbf{23.0$\times$} compared to PyTorch. Compared with PyTorch, TorchRec, HugeCTR, and OneFlow, DPIFrame can achieve state-of-the-art inference performance on GPU with speedups of \textbf{5.83$\times$}, \textbf{4.29$\times$}, \textbf{2.15$\times$}, and \textbf{2.0$\times$}, respectively.