Score
Applying transforms that map random variables to uniform(0,1) variates (componentwise or multivariate) to enable goodness-of-fit, calibration, and independence testing. The skill includes designing multivariate PITs, de-noising transformations (e.g., Poisson→Gaussian), and comparing calibration notions across discrete and continuous outcomes.
This study addresses the deviation of empirical probability integral transforms (PIT) from the theoretical uniform distribution under finite samples, a phenomenon induced by the two-stage sampling structure that invalidates conventional one-sample uniformity tests. The work systematically demonstrates that this non-uniformity arises from dependence structures and variance distortions introduced either by a reference sample or a rolling window: the former converges to a two-sample Kolmogorov–Smirnov distribution, while the latter exhibits temporal autocorrelation. Building on probability integral transform theory, empirical quantile estimation, and two-sample KS asymptotics, this paper establishes—for the first time—that empirical PIT values cannot be treated as independent uniform random variables. Leveraging these insights, the authors develop a corrected statistical inference framework specifically tailored for backtesting forecast calibration.
Conventional copula models often lack sufficient flexibility in capturing complex dependence structures. Method: This paper introduces a class of univariate W-transforms that preserve uniformity—constructed via distribution functions and piecewise strictly monotonic functions on [0,1]—ensuring transformed margins remain standard uniform. These transforms naturally induce copula-to-copula mappings, yielding W-transformed copulas. Contribution/Results: Theoretically, we derive closed-form expressions, establish conditions for density existence, and characterize rank correlations (Spearman’s ρ, Kendall’s τ), tail dependence coefficients, and symmetry properties. Methodologically, we propose an interpretable parametric family enabling independent control over central and tail dependence. Empirical results demonstrate that W-transformed copulas significantly improve fit to intricate real-world dependence patterns, particularly asymmetric tail dependence.
This study addresses goodness-of-fit testing for multivariate distributions, focusing on uniformity, normality, spherical and elliptical symmetry, and independence. It introduces a novel approach based on the decomposition of Brownian sheets. Under the null hypothesis, the problem is transformed via probability integral transforms into testing uniformity over the unit hypercube, and interactions among coordinates are effectively disentangled using a zero-margin measure decomposition. The method uniquely integrates the Gaussian process decomposition of Brownian sheets with empirical distribution functions, substantially enhancing sensitivity to joint dependence structures. Simulation studies demonstrate that the proposed test achieves power comparable to or exceeding that of current state-of-the-art methods, particularly excelling in detecting complex dependency patterns.
This paper addresses inference on general parameter transformations of cumulative distribution functions (CDFs) in the presence of nuisance parameters. We propose a unified, nonparametric, asymptotically size-controlled testing framework applicable to joint inference on one-, two-, and multi-sample CDFs. The method constructs test statistics via numerical bootstrap, obviating analytical critical value derivation and ensuring implementation simplicity. We establish theoretical guarantees of asymptotic size control and consistency. Monte Carlo simulations and empirical analyses demonstrate strong finite-sample robustness and high statistical power. Our key contribution is the first unified, nonparametric, asymptotically valid, and structure-free test for transformations of CDFs with nuisance parameters—overcoming the restrictive functional-form and parametric-structure assumptions inherent in conventional approaches.
This work addresses the longstanding limitation in conditional density estimation—namely, the absence of closed-form solutions for multivariate conditional densities under non-Gaussian assumptions. We propose a generative conditional density estimation framework grounded in copula modeling and analytic conditionalization in latent space. Methodologically, we first establish the inheritability of “conditional stability” under mixture and transformation operations, thereby extending analytically tractable conditional families to non-Gaussian, nonlinear, and cross-dimensional settings. The core components include a Gaussian Mixture Copula Model (GMCM), an explicit latent-space conditionalization mechanism, and joint copula modeling. Experiments on synthetic and real-world datasets demonstrate substantial improvements in conditional density estimation accuracy and robustness to missing data imputation. Crucially, our approach enables efficient, differentiable, and sampling-free deterministic conditional inference.
This study addresses the limitation of conventional global goodness-of-fit tests in multivariate settings, which often fail to pinpoint localized model misspecifications. To overcome this, the authors propose a local calibration test based on adaptive partitioning via Beta-trees. Departing from single-statistic global frameworks, the method evaluates whether predicted probabilities fall within finite-sample confidence intervals across data-driven subregions, enabling precise identification and visualization of model inadequacies. By leveraging k-means clustering to generate null distributions and constructing rigorous confidence intervals, the approach effectively detects local deviations in both simulated and real-world datasets, demonstrating superior performance in tasks such as selecting the number of components in mixture models.
This study addresses the limited power of traditional goodness-of-fit tests in high-dimensional settings by proposing a novel approach based on the joint distribution of multiple samples. The method employs principal component analysis for dimensionality reduction and constructs confidence sets using k-nearest neighbor estimates of high-density regions, effectively integrating information from order statistics, empirical distribution function values, and moments. It is further extended to the two-sample case. To enhance test power, the approach incorporates non-uniform transformations and probability integral transforms, with inference carried out via permutation testing. Simulation studies demonstrate that the proposed method outperforms or matches existing classical and graphical tests across a range of alternative hypotheses, substantially improving both power and applicability in high-dimensional scenarios.
This work addresses the problem of multiple hypothesis testing for multivariate Gaussian means under arbitrary covariance dependence structures. The authors propose a novel approach that integrates maximum residual descent (MRD) with a multi-stage calibration scheme. By introducing a new representation of residual statistics based on a single active precision matrix, the method achieves covariance-adaptive residualization, substantially reducing computational complexity. It replaces model-dependent thresholds with a simple multi-stage calibration rule, combining generalized stepwise critical values and precision matrix reconstruction techniques. The proposed procedure significantly lowers the normalized misclassification risk across diverse dependence structures, achieving error discovery rate control close to the nominal level, extremely low missed detection rates, near-perfect statistical power, and accurate estimation of the number of true signals.
While existing distribution testing theory has established the order of sample complexity, it lacks a precise characterization of the practical power differences among methods sharing the same asymptotic order, leaving the optimal choice ambiguous in practice. This work presents the first systematic study of sharp constants in distribution testing, focusing on uniformity and identity testing over large alphabets. By combining large-sample asymptotics under total variation distance, minimax risk bounds, and hypothesis testing theory with constant-level precision, we uncover the critical role sharp constants play in determining test power, signal-to-noise ratio, and optimal binning parameter selection. Analogous to Fisher information and Pinsker’s constant in estimation theory, our results provide a rigorous theoretical foundation for selecting testing procedures and tuning binning strategies, substantially enhancing empirical testing performance.
Traditional goodness-of-fit tests exhibit imbalanced power against diverse alternative hypotheses involving location shifts, scale changes, heavy tails, or asymmetry. This work proposes a sequential conditional calibration framework that integrates a primary test statistic—such as the Kolmogorov–Smirnov statistic—with multiple secondary statistics (e.g., variance, skewness) in a stepwise manner. The rejection regions at each stage are mutually exclusive, enabling a multiplicative decomposition of the overall Type I error rate, while the significance level of the primary test can be explicitly adjusted in subsequent stages. The method maintains high power against location alternatives and substantially enhances detection capability for scale deviations, heavy-tailed distributions, and asymmetric departures, offering an ordered decomposition of first-rejection power across stages.