🤖 AI Summary
This study addresses the bottleneck of concept drift detection in large-scale e-commerce data characterized by hundreds of millions of rows and high-dimensional features. Leveraging Apache Spark, it evaluates multi-column two-sample drift detection methods, with particular emphasis on the scalability of a distributed Maximum Mean Discrepancy (MMD) algorithm integrated with Random Fourier Features. A synthetic injection benchmark comprising 137.5 million rows is constructed, revealing a critical limitation wherein the Kolmogorov-Smirnov test fails due to statistical saturation in identifier-like columns. Experimental results demonstrate that under strong drift conditions, the proposed method achieves a Pearson correlation coefficient of 0.940, a true positive rate of 80.4%, and a false positive rate of merely 3.2%. However, its sensitivity remains constrained in weak drift scenarios.
📝 Abstract
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.