🤖 AI Summary
This work addresses the problem of multi-group multi-calibration, wherein the goal is to achieve simultaneous calibration across $k$ hierarchical statistical attributes—such as mean, variance, and value-at-risk—while minimizing sample complexity. The authors propose a randomized learning algorithm applicable to any finite family of groups and establish the first tight upper and lower bounds on sample complexity, revealing a fundamental dependence of $\varepsilon^{-(k+2)}$ on the desired calibration error $\varepsilon$ and the number $k$ of attribute levels. Leveraging tools from probability theory, statistical learning theory, and regularity conditions, they prove that for group families of polynomial size, the required sample complexity is $\widetilde{O}(\varepsilon^{-(k+2)})$. The universality of this bound is further validated across three canonical sequences of statistical attributes.
📝 Abstract
Calibration requires a predictor to be unbiased after conditioning on its own predictions. Multicalibration asks for this guarantee simultaneously across a collection of groups. Many prediction tasks ask for several related features of the same conditional outcome distribution: variance is defined relative to the mean, skewness relative to both mean and variance, and conditional value at risk relative to a quantile. We study multicalibration for a sequence of $k$ properties in which each property is identifiable once the preceding properties are fixed. This framework includes Bayes pairs but does not require the properties to arise from a single loss.
For every fixed $k\ge2$, we establish matching upper and lower sample-complexity bounds up to logarithmic factors under regularity conditions. Even with only polylogarithmically many binary groups, achieving multicalibration error $\varepsilon$ requires $\widetildeΩ(\varepsilon^{-(k+2)})$ samples. Conversely, for any finite group family $\mathcal G$, we give a randomized learner using $O(\varepsilon^{-(k+2)}+\varepsilon^{-2}\log|\mathcal G|)$ samples. Thus the sample complexity is $\widetildeΘ(\varepsilon^{-(k+2)})$ for polynomial-size group families. We instantiate the theory for three canonical examples.