Exact Comparison of Explanatory Strength of Two Dependent Predictors

📅 2026-06-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of comparing the explanatory power of two predictors for the same outcome under ill-conditioned data scenarios, such as heavy-tailed distributions or extremely sparse categories, where conventional methods often fail. The authors propose an exact nonparametric test grounded in functional exchangeability: categorical variables are handled via a pairwise swapping mechanism, while continuous variables are mapped through empirical cumulative distribution functions (ECDFs). This approach preserves marginal distributions and covariance structures without resorting to resampling that might disrupt dependencies or introduce artificial ties. The method enables exact inference in finite samples, substantially improving statistical power while rigorously controlling Type I error rates. Its robustness and effectiveness have been validated on a high-dimensional dataset of Italian compound nouns.
📝 Abstract
Comparing the relative explanatory power of two dependent predictors regarding a common target variable is a fundamental challenge across scientific disciplines. Classical asymptotic procedures, such as Vuong's closeness test or the Hotelling-Williams test, frequently collapse under pathological data conditions, including heavy-tailed distributions and extreme categorical sparsity. To bypass these limitations, practitioners often turn to non-parametric resampling. However, naive permutation tests destroy the natural covariance structure of dependent predictors, while the paired bootstrap evaluating variance around the alternative hypothesis and introducing artificial ties suffers from metric space compression and categorical omission, rendering it highly unreliable in finite samples. In this paper, we introduce the Paired Swap Permutation Test, a novel and exact non-parametric methodology. Grounded in the principle of functional exchangeability under the null hypothesis, our algorithm utilizes a symmetric within-subject swapping mechanism for categorical data, and introduces an Empirical Cumulative Distribution Function (ECDF) mapping step for continuous domains. This copula-based transposition perfectly preserves marginal densities and the empirical support without introducing resampling ties. Through extensive Monte Carlo simulations, we demonstrate that the proposed test strictly maintains the nominal significance level and maximizes statistical power under conditions where standard methods become catastrophically liberal or pathologically conservative. Finally, we apply the framework to a high-dimensional linguistic dataset of Italian noun-noun compounds, proving its capacity to deliver robust, exact inference in environments where conventional analytical methods inherently fail.
Problem

Research questions and friction points this paper is trying to address.

explanatory strength
dependent predictors
non-parametric comparison
categorical sparsity
heavy-tailed distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Paired Swap Permutation Test
functional exchangeability
ECDF mapping
non-parametric inference
dependent predictors
💼 Related Jobs
No related jobs found.