distribution-free testing

Design and analyze hypothesis tests and testing procedures that do not rely on parametric or specific distributional assumptions; build test statistics and decision rules and prove formal guarantees such as impossibility and achievability results, distribution-free error bounds, and sample-complexity or minimax-style performance guarantees.

distribution-freetesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the problem of diminished statistical power in nonparametric tests arising from excessive dependence between test statistics and auxiliary statistics. To resolve this, we propose a novel framework grounded in statistical independence principles. Methodologically, we reformulate hypothesis testing via a relativity principle and establish asymptotic independence between test and auxiliary statistics—yielding a Basu-type theoretical guarantee—while preserving distributional invariance under both null and alternative hypotheses. Integrating decision-theoretic criteria with explicit independence constraints, our approach systematically enhances classical tests, including Shapiro–Wilk, Anderson–Darling, Kolmogorov–Smirnov, and symmetry-center tests. Extensive simulations demonstrate that the proposed methods significantly improve power over conventional approaches while retaining robustness and computational efficiency, making them broadly applicable across diverse nonparametric settings.

Develops test statistics via ancillary independence for improved accuracyEnhances classical nonparametric tests through principled methodologyMinimizes dependence between test statistics and ancillary structures

StatWhy: Formal Verification Tool for Statistical Hypothesis Testing Programs

May 25, 2024
YK
Yusuke Kawamoto
🏛️ AIST | PRESTO | JST | University of Tsukuba | Kyoto University

Misuse of statistical hypothesis tests severely undermines scientific reliability. This paper proposes a formal verification methodology for statistical programs: preconditions—such as normality, independence, and homoscedasticity—are explicitly encoded as logical assertions in source code; static verification is then performed on OCaml implementations using the Why3 platform to automatically detect missing or conflicting assumptions. The approach innovatively integrates contract-based programming with formal verification, distinguishing between formalizable preconditions (amenable to automated checking) and non-formalizable ones (requiring expert judgment), thereby establishing a human-in-the-loop verification paradigm. Evaluated on canonical statistical tests—including Student’s *t*-test and ANOVA—the method successfully identifies widespread misuses, such as applying the *t*-test to non-normal data or neglecting homoscedasticity checks. Results demonstrate significant improvements in the correctness, auditability, and reproducibility of statistical software.

Automatically check requirements for statistical methods in codeFormally verify correctness of statistical hypothesis testing programsPrevent common errors in statistical program implementation

This study addresses the risk of misleading inferences and Type III errors—rejecting a hypothesis that is neither the true null nor the true alternative—in order-specified hypothesis testing, which often arises from misspecified constraints. To mitigate this issue, the paper introduces, for the first time, a “safe hypothesis testing” framework that incorporates a validity certificate, a pre-test mechanism assessing the compatibility between the imposed constraints and the observed data. Rejection of the null hypothesis is permitted only when the constraints are deemed reasonable, thereby preventing systematic errors. Integrating order-restricted inference theory, asymptotic analysis, and pre-testing, the proposed method controls Type III error rates while maintaining statistical power comparable to conventional approaches, thus achieving both reliability and efficiency. Extensive simulations and empirical analyses demonstrate its superior performance.

hypothesis testingmisspecified constraintsorder-restricted inference

Distribution Testing in the Presence of Arbitrarily Dominant Noise with Verification Queries

Sep 21, 2025
HB
Hadley Black
🏛️ University of California, San Diego

This work studies efficient distribution testing under two challenging conditions: (i) only a small number of relevant samples are accessible, and (ii) the observed data is arbitrarily corrupted by dominant noise. To address this, we propose a novel “verification query” model—a query mechanism designed to statistically distinguish the target distribution from noise without direct access to clean samples. Our theoretical analysis establishes a smooth trade-off between sample complexity and query complexity. We derive the first tight upper and lower bounds for uniformity, identity, and closeness testing in this setting. Crucially, when the probability density function (PDF) of the mixture distribution is available, our approach breaks classical lower bounds, significantly reducing the required number of queries while provably robust against adaptive adversarial noise.

Developing verification queries to identify noise-contaminated samplesEstablishing trade-offs between sample size and query complexityTesting distributions with limited access to relevant data samples

This paper investigates the sample complexity of Classification Accuracy Testing (CAT) for nonparametric hypothesis testing. We formulate and systematically analyze CAT across three canonical nonparametric settings: discrete distributions, $d$-dimensional distributions with Hölder-continuous densities, and the Gaussian sequence model. CAT trains a binary classifier on synthetic samples and constructs the test statistic from its empirical classification accuracy. We establish, for the first time, that CAT achieves (near) minimax-optimal sample complexity under total variation distance $varepsilon$-separation and type-II error probability $delta$, thereby closing the high-probability complexity gap for likelihood-free testing. Notably, in the discrete two-sample setting, CAT exactly recovers the known minimax-optimal rate. These results provide a theoretically complete foundation for density-ratio-free nonparametric testing.

Closing high probability sample complexity for likelihood-free hypothesis testingEstablishing minimax optimality in goodness-of-fit and two-sample testingOptimizing sample complexity for classification-based hypothesis testing methods

Latest Papers

What's happening recently
View more

This study addresses the challenge of performing effective sequential hypothesis testing on simulated data lacking explicit density functions, a scenario where existing methods often fall short. To overcome this limitation, this work proposes a density-free, efficient solution by integrating martingale theory, e-value statistics, and sequential analysis techniques into a unified e-test martingale framework. The primary contribution lies in enabling anytime-valid sequential testing that simultaneously achieves growth optimality and rigorous error control. Theoretically, the proposed method guarantees strict control of Type I error at arbitrary stopping times, ensures geometric decay of Type II error, and yields asymptotic power approaching unity. Consequently, this framework establishes a reliable anytime-valid paradigm for simulator-driven statistical inference.

anytime-validdensity-freee-test martingale

This work addresses the challenge of distribution identity testing in high-dimensional or continuous domains, where traditional total variation distance becomes ineffective. The authors propose a fooling distance framework based on bounded discriminator classes, integrating integral probability metrics with Boolean function classes to achieve sample-efficient testing algorithms. By unifying three directions—testable learning, verification of learning, and structured distribution testing—the study extends the Ak-testing framework and establishes theoretical connections among them. Technically, the approach combines Rademacher complexity analysis, membership query mechanisms, and polynomial density modeling over the hypercube, yielding testable proper learners for halfspaces and decision trees, deriving verification lower bounds, and designing efficient identity testers for decision tree distributions and low-degree polynomial densities.

bounded distinguishersdistribution testingfooling distance

This work addresses the problem of determining the minimum sample complexity required to test halfspaces in a distribution-free setting where only random samples are accessible. By combining probabilistic analysis with information-theoretic techniques, the authors constructively design a one-sided tester and establish a matching lower bound via an adversarial argument, thereby proving for the first time a tight bound of Θ(n/ε) on the sample complexity. This result demonstrates that testing and learning share the same fundamental efficiency limit in this model, and that one-sided testers are already optimal—bilateral testers offer no additional advantage. Consequently, the findings refine the conventional testing-to-learning reduction framework by showing that optimal testing can be achieved without resorting to learning-based approaches.

distribution-freehalfspace testinglower bound

This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.

familywise errorhypothesis configurationmultiple hypotheses

This study addresses the problem of distribution equivalence testing under the non-adaptive conditional sampling model, aiming to determine whether two unknown distributions are identical. Methodologically, by integrating techniques from distribution testing theory and information-theoretic lower bound analysis, it systematically derives theoretical bounds on query complexity. The primary contributions are threefold: first, it proves that uniformity, identity, and equivalence testing all exhibit a query complexity of Θ̃(log n) within this model; second, it establishes that the necessary and sufficient number of queries for equivalence testing is Θ̃(log n/ε²). These findings yield the first tight bound for this problem, thereby providing a complete theoretical characterization of distribution testing under the conditional sampling framework.

distribution testingequivalence testingnon-adaptive conditional samples

Hot Scholars

SC

Sitan Chen

Assistant Professor of Computer Science, Harvard University
theoretical computer sciencegenerative modelingquantum informationmathematics of data science
RO

Ryan O'Donnell

Professor, Computer Science Department, Carnegie Mellon
Theoretical Computer Science
SL

Sihan Liu

University of California San Diego
Theoretical Machine LearningDistribution Testing
ID

Ilias Diakonikolas

University of Wisconsin-Madison
theoretical computer sciencealgorithmic statisticsmachine learningprobability theory