SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data

📅 2026-05-01
📈 Citations: 0
Influential: 0
📄 PDF

career value

220K/year
🤖 AI Summary
This work addresses the tension between logical data partitioning in production environments and efficient GPU utilization by proposing a streaming GPU encoding system. The system introduces a SuperBatch mechanism to uniformly process heterogeneous partitioned data, integrating a precise throughput prediction model, a streaming dual-threshold strategy bounded by memory safety margins, and an adaptability decision framework (φ/CV) to simultaneously achieve high throughput, bounded memory usage, low latency, and strong fault tolerance. By incorporating zero-copy Arrow serialization and an asynchronous I/O pipeline, the system attains a throughput of 26,413 samples per second on a 10-million-text dataset while consuming only 2.6 GB of memory—1/12.6 of the baseline—and reduces first-output latency by 68×, all while maintaining a 7% throughput advantage over the baseline.
📝 Abstract
We present SURGE, a streaming GPU encoding system deployed in production to generate embeddings for over 800 million texts across 40,000 logical partitions. Production embedding pipelines face a tension between logical data partitioning and efficient GPU utilization: processing each partition independently incurs $P$ inter-process communication (IPC) calls whose overhead limits throughput for compute-light models. Our contributions are analytical: (i) a cost model (Theorem 1) predicting throughput within 2% across three encoders spanning a 15$\times$ parameter range; (ii) a memory-safety bound (Lemma 3) enabling a streaming two-threshold policy with peak memory $O(B_{\min} + n_{\max})$ rather than $O(N)$; and (iii) a $φ$/CV decision framework characterizing when the pattern applies beyond our workload. The naive fix of batching at fixed size requires $O(N)$ peak memory (32.7 GB at 10M texts; infeasible beyond ~60M on 192 GB nodes), produces no output until all encoding completes, and offers no fault tolerance. SURGE achieves the same throughput with $O(B_{\min} + n_{\max})$ bounded memory (2.6 GB), 68$\times$ faster time-to-first-output, and crash recovery at SuperBatch granularity. On 10M texts with 4 NVIDIA L4 GPUs, SURGE delivers 26,413 texts/s -- matching fixed-batch throughput while using 12.6$\times$ less memory. We validate on bge-base (109M, $d$=768, error 1.3%) and across log-normal $σ$ in {1.0, 1.72, 2.5} (speedup invariant within $\pm$3%), and compare against a partition-batched baseline (PB-PBP-LB), against which SURGE retains a 7% throughput edge and 2.5$\times$ faster TTFO. Complementary engineering -- zero-copy Arrow serialization (22-25$\times$ speedup) and async I/O pipelining (up to 93% benefit) -- realizes the design but is not the contribution.
Problem

Research questions and friction points this paper is trying to address.

GPU encoding
heterogeneous partitioned data
memory efficiency
throughput optimization
embedding pipelines
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPU encoding
memory-efficient streaming
cost model
fault tolerance
heterogeneous partitioned data
S
Shashank Kapadia
Walmart Inc.
D
Deep Narayan Mishra
Walmart Inc.
S
Sujal Reddy Alugubelli
Walmart Inc.
A
Ajay Kumar
Walmart Inc.
S
Swapnil Yadav
Walmart Inc.
R
Rishi Bhatia
Walmart Inc.