signature two-sample testing

Designs and implements nonparametric two‑sample hypothesis tests for distributions over path- or stream-valued data by building test statistics from path signature features or from the signature kernel; this includes constructing RKHS embeddings or feature summaries, computing significance (e.g., permutation, bootstrap, or asymptotic) and evaluating the test's power to detect distributional differences between samples.

signaturetwo-sampletesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the two-sample testing problem without distributional assumptions. It proposes PReLU-TST, a novel approach based on integral probability metrics (IPMs), which constructs a nonparametric test statistic using a parametrized discriminator consisting of only a single neuron. This design retains the flexibility of nonparametric methods while substantially improving computational efficiency. The proposed test is proven to be consistent and asymptotically equivalent to classical nonparametric IPM-based tests. Empirical evaluations demonstrate that PReLU-TST achieves higher or at least comparable finite-sample testing power against existing methods across a range of synthetic and real-world datasets.

distributional differenceintegral probability metricnonparametric test

Distance and Kernel-Based Measures for Global and Local Two-Sample Conditional Distribution Testing

Oct 15, 2022
JY
Jian Yan
🏛️ Cornell University | Texas A&M University

This paper addresses the critical yet underexplored problem of testing equivalence between two conditional distributions—a fundamental task in transfer learning and causal inference. We propose the first unified framework for both global and local two-sample conditional distribution testing. Our method introduces: (1) novel distance and kernel-based metrics that characterize conditional distribution homogeneity; (2) an estimation theory grounded in conditional U-statistics, enabling integrated modeling for both global and local tests; and (3) a principled combination of RKHS embeddings and localized bootstrap resampling, yielding convergence rates and asymptotic null/alternative distributions of the estimators. Theoretical analysis guarantees strong statistical power, while empirical evaluations on synthetic and real-world datasets demonstrate high detection accuracy and robustness.

Developing distance and kernel-based measures for distribution homogeneityProposing consistent estimators and hypothesis tests for conditional distributionsTesting equality of two conditional distributions globally and locally

A Permutation-free Kernel Two-Sample Test

Nov 27, 2022
SS
S. Shekhar
🏛️ Carnegie Mellon University | Yonsei University

Traditional kernel Maximum Mean Discrepancy (MMD) two-sample tests rely on permutation to determine critical thresholds, ensuring finite-sample validity but incurring an O(n²) computational cost per permutation—prohibitively expensive for large samples. This paper proposes the cross-MMD test statistic: by splitting samples to construct a U-statistic, and combining studentization with a Gaussian kernel, it yields the first kernel MMD test that requires no permutations. The method achieves asymptotic normality with a single O(n²) computation, while preserving finite-sample validity, statistical consistency, and minimax optimal detection rates under local alternatives. Theoretically and empirically, cross-MMD accelerates testing by over an order of magnitude compared to permutation-based approaches on large samples, with only a marginal loss in power, and maintains strong consistency against any fixed distributional discrepancy.

Achieves asymptotic normality and minimax optimal powerAvoids computationally expensive permutation methodsProposes cross-MMD for efficient two-sample testing

Compress Then Test: Powerful Kernel Testing in Near-linear Time

Jan 14, 2023
CD
Carles Domingo-Enrich
🏛️ NYU | Harvard University | MIT | Microsoft Research

Two-sample kernel testing has long faced a trade-off between statistical power and computational efficiency: exact tests require $O(n^2)$ time, while existing acceleration methods substantially degrade detection power. This paper proposes the “Compress-Then-Test” (CTT) framework, which constructs high-fidelity coresets to enable near-linear-time testing ($O(n log n)$) while provably preserving the optimal detection boundary of quadratic-time kernel tests under sub-exponential distributions—the first method to achieve this guarantee. Theoretically, we prove asymptotic equivalence of the compressed test statistic, enabling rigorous, fast permutation testing. CTT integrates sample compression, low-rank kernel approximation, adaptive kernel selection, and an improved permutation strategy. Experiments on synthetic and real-world datasets demonstrate that CTT accelerates state-of-the-art approximate MMD methods by 20–200× without sacrificing statistical power.

Improves efficiency without sacrificing detection accuracyMaintains high statistical power via sample compressionReduces kernel test runtime from quadratic to near-linear

Practical Kernel Tests of Conditional Independence

Feb 20, 2024
RP
Roman Pogodin
🏛️ McGill University | Mila | Gatsby Computational Neuroscience Unit | University College London | University of British Columbia | Amii

Addressing the dual challenges of inflated Type I error rates (loss of test-level control) and low statistical power in conditional independence testing, this paper proposes a data-efficient kernel-based testing framework. The method employs kernel ridge regression and introduces, for the first time in this setting, three principled bias-correction strategies: data splitting, auxiliary data utilization, and restriction to simplified function classes—ensuring rigorous asymptotic and finite-sample control of the significance level. Theoretically, the approach guarantees convergence of the Type I error rate to the nominal significance level while enhancing detection power for complex dependency structures. Extensive experiments on diverse synthetic and real-world datasets demonstrate that the proposed method achieves precise Type I error control and substantially outperforms state-of-the-art competitors—including KCIT and RCIT—in statistical power, with improved robustness and reliability.

Addressing bias in test statistics from kernel ridge regression methodsDeveloping kernel-based tests for conditional independence with accurate false positive controlImproving test level accuracy while maintaining competitive statistical power

Latest Papers

What's happening recently
View more

This work proposes a nonparametric method to assess the statistical significance of signal features—such as peaks and plateaus—in data and to detect multimodal structures in inter-event spacing distributions. The approach leverages run theory, employing a Markov chain recursion to precisely characterize the distribution of the longest runs. It integrates permutation testing with a bootstrap procedure tailored for continuous data, enabling a unified evaluation of both high- and low-intensity signal features. The key innovation lies in the first principled synthesis of run-length analysis, permutation tests, and continuous-data bootstrapping, which collectively facilitate accurate detection and localization of multimodal patterns without requiring parametric distributional assumptions, thereby effectively identifying salient morphological features in complex datasets.

bootstrap testmulti-modalitynon-parametric

Novelty detection on path space

Dec 02, 2025
IG
Ioannis Gasteratos
🏛️ TU Berlin | Imperial College London | University of Oxford

This paper addresses the lack of rigorous statistical guarantees in novelty detection on path space. Methodologically, it introduces the first nonparametric hypothesis testing framework based on signature statistics: it constructs a smooth CVaR surrogate objective using the shuffle product identity of path signatures and leverages transport cost inequalities to control Type-I error for non-Gaussian processes—including laws of rough differential equations (RDEs). A novel SVM-based algorithm optimizes this objective, enabling computable estimates of quantiles and p-values. Contributions include: (i) the first incorporation of transport inequalities into path-space novelty detection; (ii) exact false positive rate control without Gaussianity assumptions; (iii) theoretical lower bounds on Type-II error and a general power bound under absolutely continuous alternative hypotheses. Empirical validation on synthetic anomalous diffusion data and real molecular biology datasets confirms statistical power and robustness.

Derive tail bounds for false positive rates beyond Gaussian measuresDevelop hypothesis testing for novelty detection on path spaceEstablish error bounds and evaluate signature-based test statistics

This study addresses the problem of constructing a nonparametric sequential hypothesis test with unit power when only historical offline data are available to implicitly define a null hypothesis and multiple alternative distributions. To this end, the work introduces— for the first time—a multiclass classifier into the sequential testing framework, proposing a procedure that controls the significance level at α and almost surely identifies the true underlying distribution. Under mild separability conditions, the method is theoretically shown to possess a tight upper bound on its stopping time, thereby achieving optimal stopping performance. The proposed approach naturally accommodates both distribution identification and scenarios involving train-test distribution mismatch. Its effectiveness is validated through comprehensive experiments on both synthetic and real-world datasets.

classifier-basednonparametricoffline data

This work addresses the limited power of nonparametric two-sample tests in high-dimensional or complex distributional settings by proposing the spectrally truncated normalized Maximum Mean Discrepancy (st-nMMD). Built upon embeddings in a reproducing kernel Hilbert space, st-nMMD integrates covariance operator normalization with spectral truncation regularization to substantially enhance test power. The paper establishes, for the first time, a non-asymptotic exponential upper bound for st-nMMD under the null hypothesis, introduces an adaptive hyperparameter tuning algorithm that avoids data splitting, and provides explicit non-asymptotic quantile estimates. Empirical results demonstrate that the method maintains proper Type I error control while achieving superior statistical power and stability under the alternative hypothesis, significantly outperforming existing kernel-based two-sample tests.

kernel methodsnon-asymptotic analysisnormalized MMD

This work proposes a nonparametric kernel-based approach for inference in multivariate or functional time series, addressing problems such as goodness-of-fit testing, change-point detection in marginal distributions, and independence testing. The method avoids both resampling and bandwidth selection by embedding the data into a reproducing kernel Hilbert space (RKHS) and constructing test statistics through sample splitting, projection, and self-normalization. Leveraging a novel conditioning technique, the authors establish that the resulting test statistic admits a pivotal asymptotic null distribution under strong mixing conditions and analyze its power against local alternatives. The proposed procedure achieves high finite-sample accuracy while substantially improving computational efficiency, outperforming existing resampling-based methods.

change pointgoodness-of-fitnonparametric inference

Hot Scholars

LF

Livio Finos

Full Professor of Statistics, University of Padova
multivariate analysispermutation testsbiostatisticspsychometrics
WZ

Wenjie Zhao

University of Texas at Dallas
computer vision
LH

Liam Hodgkinson

University of Melbourne
probabilistic machine learningdeep learning theory
PV

Petar Veličković

Senior Staff Research Scientist, Google DeepMind | Affiliated Lecturer, University of Cambridge
Geometric Deep LearningGraph Neural NetworksCategorical Deep LearningAlgorithmic Reasoning