A Quick and Exact Method for Distributed Quantile Computation

📅 2025-11-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Precise quantile computation in large-scale distributed environments faces fundamental challenges: it traditionally requires global sorting, incurs prohibitive communication overhead, and struggles to balance accuracy with efficiency. To address this, we propose GK Select—an algorithm that pioneers the use of the Greenwald-Khanna (GK) sketch for *exact* quantile computation. GK Select operates by extracting candidate values within the GK error bound, performing partition-wise linear scans, and applying a tree-based reduction—achieving exact results without full data shuffling in a constant number of communication rounds. Theoretically, its time complexity matches that of the GK sketch, while its space complexity is O(1/ε). Experiments on a 30-core AWS EMR cluster demonstrate that GK Select outperforms Spark’s built-in global-sorting approach by 10.5× in throughput and achieves latency comparable to approximate methods—thereby substantially overcoming the scalability bottleneck of exact quantile computation.

Technology Category

Search and Optimization: Distributed SearchData Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsMachine Learning: Calibration & Uncertainty Quantification

Application Category

Graph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphsSecurity and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
Quantile computation is a core primitive in large-scale data analytics. In Spark, practitioners typically rely on the Greenwald-Khanna (GK) Sketch, an approximate method. When exact quantiles are required, the default option is an expensive global sort. We present GK Select, an exact Spark algorithm that avoids full-data shuffles and completes in a constant number of actions. GK Select leverages GK Sketch to identify a near-target pivot, extracts all values within the error bound around this pivot in each partition in linear time, and then tree-reduces the resulting candidate sets. We show analytically that GK Select matches the executor-side time complexity of GK Sketch while returning the exact quantile. Empirically, GK Select achieves sketch-level latency and outperforms Spark's full sort by approximately 10.5x on 10^9 values across 120 partitions on a 30-core AWS EMR cluster.
Problem

Research questions and friction points this paper is trying to address.

Exact quantile computation avoids expensive global sorting
Distributed algorithm eliminates full-data shuffle operations
Maintains sketch-level latency while ensuring precise results
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses GK Sketch to identify near-target pivot
Extracts candidate values within error bounds linearly
Tree-reduces candidate sets for exact quantile computation
I
Ivan Cao
Computer and Information Science, University of Mississippi, Oxford, MS, USA
J
Jaromir J. Saloni
Computer and Information Science, University of Mississippi, Oxford, MS, USA
D
David A. G. Harrison
Computer and Information Science, University of Mississippi, Oxford, MS, USA