Why Empirical p-Values Are Not Uniform: Reference Samples, Dependence, and PIT Backtesting

📅 2026-05-15
📈 Citations: 0
Influential: 0
📄 PDF

career value

180K/year
🤖 AI Summary
This study addresses the deviation of empirical probability integral transforms (PIT) from the theoretical uniform distribution under finite samples, a phenomenon induced by the two-stage sampling structure that invalidates conventional one-sample uniformity tests. The work systematically demonstrates that this non-uniformity arises from dependence structures and variance distortions introduced either by a reference sample or a rolling window: the former converges to a two-sample Kolmogorov–Smirnov distribution, while the latter exhibits temporal autocorrelation. Building on probability integral transform theory, empirical quantile estimation, and two-sample KS asymptotics, this paper establishes—for the first time—that empirical PIT values cannot be treated as independent uniform random variables. Leveraging these insights, the authors develop a corrected statistical inference framework specifically tailored for backtesting forecast calibration.
📝 Abstract
Probability integral transforms (PITs) and empirical $p$-values are widely used to assess the calibration of predictive distributions. While exact PIT values are uniformly distributed under correct model specification, practical implementations rely on empirical estimates constructed from finite samples. We show that this estimation step fundamentally alters the statistical structure of the problem. In particular, common-sample and rolling-window implementations introduce dependence and variance distortions that invalidate classical one-sample uniformity tests. When empirical percentiles are conditioned on a shared reference sample, the resulting statistics converge towards a two-sample Kolmogorov--Smirnov regime, while rolling windows induce autocorrelation and variance suppression. Our findings indicate that treating empirical percentiles as independent uniform draws can distort statistical inference and that backtesting procedures based on PITs require revised calibration methods accounting for the underlying two-stage sampling structure.
Problem

Research questions and friction points this paper is trying to address.

empirical p-values
probability integral transform
uniformity test
dependence
backtesting
Innovation

Methods, ideas, or system contributions that make the work stand out.

empirical p-values
probability integral transform
dependence structure
two-sample Kolmogorov–Smirnov
backtesting calibration
🔎 Similar Papers
No similar papers found.