🤖 AI Summary
Precise quantile computation in large-scale distributed environments faces fundamental challenges: it traditionally requires global sorting, incurs prohibitive communication overhead, and struggles to balance accuracy with efficiency. To address this, we propose GK Select—an algorithm that pioneers the use of the Greenwald-Khanna (GK) sketch for *exact* quantile computation. GK Select operates by extracting candidate values within the GK error bound, performing partition-wise linear scans, and applying a tree-based reduction—achieving exact results without full data shuffling in a constant number of communication rounds. Theoretically, its time complexity matches that of the GK sketch, while its space complexity is O(1/ε). Experiments on a 30-core AWS EMR cluster demonstrate that GK Select outperforms Spark’s built-in global-sorting approach by 10.5× in throughput and achieves latency comparable to approximate methods—thereby substantially overcoming the scalability bottleneck of exact quantile computation.
📝 Abstract
Quantile computation is a core primitive in large-scale data analytics. In Spark, practitioners typically rely on the Greenwald-Khanna (GK) Sketch, an approximate method. When exact quantiles are required, the default option is an expensive global sort. We present GK Select, an exact Spark algorithm that avoids full-data shuffles and completes in a constant number of actions. GK Select leverages GK Sketch to identify a near-target pivot, extracts all values within the error bound around this pivot in each partition in linear time, and then tree-reduces the resulting candidate sets. We show analytically that GK Select matches the executor-side time complexity of GK Sketch while returning the exact quantile. Empirically, GK Select achieves sketch-level latency and outperforms Spark's full sort by approximately 10.5x on 10^9 values across 120 partitions on a 30-core AWS EMR cluster.