Compress Then Test: Powerful Kernel Testing in Near-linear Time

📅 2023-01-14
🏛️ International Conference on Artificial Intelligence and Statistics
📈 Citations: 9
✨ Influential: 1
📄 PDF
🤖 AI Summary
Two-sample kernel testing has long faced a trade-off between statistical power and computational efficiency: exact tests require $O(n^2)$ time, while existing acceleration methods substantially degrade detection power. This paper proposes the “Compress-Then-Test” (CTT) framework, which constructs high-fidelity coresets to enable near-linear-time testing ($O(n log n)$) while provably preserving the optimal detection boundary of quadratic-time kernel tests under sub-exponential distributions—the first method to achieve this guarantee. Theoretically, we prove asymptotic equivalence of the compressed test statistic, enabling rigorous, fast permutation testing. CTT integrates sample compression, low-rank kernel approximation, adaptive kernel selection, and an improved permutation strategy. Experiments on synthetic and real-world datasets demonstrate that CTT accelerates state-of-the-art approximate MMD methods by 20–200× without sacrificing statistical power.
📝 Abstract
Kernel two-sample testing provides a powerful framework for distinguishing any pair of distributions based on $n$ sample points. However, existing kernel tests either run in $n^2$ time or sacrifice undue power to improve runtime. To address these shortcomings, we introduce Compress Then Test (CTT), a new framework for high-powered kernel testing based on sample compression. CTT cheaply approximates an expensive test by compressing each $n$ point sample into a small but provably high-fidelity coreset. For standard kernels and subexponential distributions, CTT inherits the statistical behavior of a quadratic-time test -- recovering the same optimal detection boundary -- while running in near-linear time. We couple these advances with cheaper permutation testing, justified by new power analyses; improved time-vs.-quality guarantees for low-rank approximation; and a fast aggregation procedure for identifying especially discriminating kernels. In our experiments with real and simulated data, CTT and its extensions provide 20--200x speed-ups over state-of-the-art approximate MMD tests with no loss of power.
Problem

Research questions and friction points this paper is trying to address.

Reduces kernel test runtime from quadratic to near-linear
Maintains high statistical power via sample compression
Improves efficiency without sacrificing detection accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compress samples into high-fidelity coresets
Achieves near-linear time with optimal power
Combines compression with cheaper permutation testing
🔎 Similar Papers