Optimal Rates for Learning with Monotone Adversaries

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the optimal error rates of statistical learning under monotonic adversarial perturbations, where an adversary appends to the original sample additional points that depend on it and are correctly labeled. For non-exchangeable distributions, the authors develop a unified minimax framework that characterizes learnability through both the VC dimension and the Littlestone dimension. Leveraging minimax analysis, leave-one-out arguments, and inclusion graphs, they establish that when the VC dimension $d \geq 2$, the minimax expected error is $\Theta\left(\frac{d}{n}\log\frac{n}{d}\right)$, whereas for $d = 1$, it improves to $\Theta(1/n)$. These results demonstrate that, except in the trivial case $d=1$, monotonic adversarial interference inherently incurs a logarithmic penalty, thereby invalidating the applicability of classical online-to-batch conversion rates in this setting.
📝 Abstract
A monotone adversary observes an i.i.d. labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. The learner sees a uniform shuffle of the combined sample and is scored on the original distribution. Every example is correctly labeled, but the insertions depend on the clean sample, so the combined sample is not exchangeable. Larsen, Pabbaraju, and Shetty, who introduced this model, showed that empirical risk minimization attains expected error $O((d/n)\log(n/d))$ for classes of VC dimension $d$, and that every known optimal learner can be pushed away from the $Θ(d/n)$ rate, optimal for PAC learning. They asked whether the extra logarithm is an artifact of those particular algorithms or an inherent consequence of the lack of exchangeability. We show that this additional cost is inherent beyond VC dimension one. In the worst case over classes of VC dimension $d$ and over known finite insertion budgets, the minimax expected error is $Θ(1/n)$ at $d=1$ and $Θ((d/n)\log(n/d))$ for $d\geq 2$. The same rates hold with Littlestone dimension $d_{\mathrm L}$ in place of $d$, so the clean online-to-batch rate $O(d_{\mathrm L}/n)$ is unattainable as well. Thus, somewhat counterintuitively, adding correctly labeled examples can make learning harder by a logarithmic factor, even for classes that admit finite mistake bounds in online learning. The dimension-one upper bound is achieved by a simple improper learner whose analysis adapts the leave-one-out argument underlying the one-inclusion graph. All of our lower bounds are elementary and come from a single construction: an explicit class and prior on which two target hypothesis, which differ a point of nonnegligible mass, produce the same sample.
Problem

Research questions and friction points this paper is trying to address.

monotone adversaries
VC dimension
minimax error rate
non-exchangeability
PAC learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

monotone adversary
minimax rate
VC dimension
non-exchangeability
online-to-batch conversion
🔎 Similar Papers
No similar papers found.