A Fast and Effective Large-Scale Two-Sample Test Based on Kernels

📅 2021-10-07
📈 Citations: 3
Influential: 0
📄 PDF

career value

249K/year
🤖 AI Summary
To address the high computational cost, low statistical power, and bandwidth sensitivity of kernel two-sample tests on high-dimensional, large-scale data, this paper proposes a parameter-free robust kernel test. The method avoids bandwidth selection entirely while ensuring reliability and high power. Its core contributions are threefold: (1) a novel test statistic designed via theoretical analysis to eliminate power loss from data splitting; (2) a non-asymptotic significance control mechanism guaranteeing validity under finite samples; and (3) inherent suitability for high-dimensional settings, delivering uniformly high power across diverse alternative hypotheses. Experiments on synthetic and real-world datasets demonstrate that the proposed method achieves 10–100× speedup over MMD and state-of-the-art large-scale kernel tests, with average power gains of 15%–40%, all without any bandwidth tuning.
📝 Abstract
Kernel two-sample tests have been widely used and the development of efficient methods for high-dimensional large-scale data is gaining more and more attention as we are entering the big data era. However, existing methods, such as the maximum mean discrepancy (MMD) and recently proposed kernel-based tests for large-scale data, are computationally intensive to implement and/or ineffective for some common alternatives for high-dimensional data. In this paper, we propose a new test that exhibits high power for a wide range of alternatives. Moreover, the new test is more robust to high dimensions than existing methods and does not require optimization procedures for the choice of kernel bandwidth and other parameters by data splitting. Numerical studies show that the new approach performs well in both synthetic and real world data.
Problem

Research questions and friction points this paper is trying to address.

Develops efficient kernel test for large-scale data
Addresses computational intensity of existing two-sample tests
Improves power and robustness in high-dimensional settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

New kernel test for large-scale two-sample data
Robust high-dimensional performance without parameter optimization
Computationally efficient with wide alternative power range
🔎 Similar Papers