🤖 AI Summary
This work addresses the failure of standard exchangeability—and consequent unreliability of distribution-free inference—under hierarchical data structures (e.g., grouped or repeated-measures designs). We introduce *hierarchical exchangeability*, the first formal theoretical foundation for distribution-free inference in non-i.i.d. hierarchical settings. Methodologically, we extend conformal prediction and the jackknife+ framework to hierarchical structures and propose a second-moment coverage mechanism, strengthening guarantees from marginal coverage to *conditional second-moment coverage*. Experiments demonstrate that our approach substantially reduces conditional miscoverage rates; under model misspecification, prediction interval width increases only marginally, while under correct model specification, the overhead is negligible. Our core contribution is the first provably reliable, distribution-free inference framework tailored specifically for hierarchical data.
📝 Abstract
This paper studies distribution-free inference in settings where the data set has a hierarchical structure -- for example, groups of observations, or repeated measurements. In such settings, standard notions of exchangeability may not hold. To address this challenge, a hierarchical form of exchangeability is derived, facilitating extensions of distribution-free methods, including conformal prediction and jackknife+. While the standard theoretical guarantee obtained by the conformal prediction framework is a marginal predictive coverage guarantee, in the special case of independent repeated measurements, it is possible to achieve a stronger form of coverage -- the"second-moment coverage"property -- to provide better control of conditional miscoverage rates, and distribution-free prediction sets that achieve this property are constructed. Simulations illustrate that this guarantee indeed leads to uniformly small conditional miscoverage rates. Empirically, this stronger guarantee comes at the cost of a larger width of the prediction set in scenarios where the fitted model is poorly calibrated, but this cost is very mild in cases where the fitted model is accurate.