🤖 AI Summary
This study addresses whether a separation exists between the learning and prediction sample complexities for Littlestone classes in differentially private online learning. By integrating differential privacy theory, Littlestone dimension analysis, and combinatorial optimization techniques, the authors establish both lower bounds for private online learning and upper bounds for prediction. This work makes the first contribution to filling the gap of non-trivial lower bounds within specific parameter regimes. It rigorously proves that learning requires $\Omega(\log T^{2/3})$ mistakes, whereas prediction necessitates only $O(\log \log T^2)$ mistakes. These results reveal a significant complexity divergence as time progresses, formally establishing the separation between learning and prediction sample complexities in differentially private online settings.
📝 Abstract
We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs to make predictions without releasing a hypothesis. Using a novel lower bound for private online learning and an upper bound for private prediction, we show that the sample complexity of these two problems are separated by a factor that grows with the time horizon for every class of finite Littlestone dimension $d$. First, we prove that every $\br{\epsilon,\delta}$-private online learner has a deterministic realisable stream of length $T$ on which the mistake bound is at least $\bE\bs{M_T}=\Om{\frac d\epsilon \log\br{ T}^{2/3}}$. In particular, this is the first non-trivial lower in the range $1/T<\delta<1/\log T)$ left open in earlier works[SR22,DSS24,LWY24]. Second, we prove that for every class of of Littlestone dimension $d$, there exists an $(\epsilon,\delta)$-jointly private predictor with at most $2^{2^{cd^2}}\epsilon^{-2}\log^2\br{2/\br{\epsilon\delta}}$ expected mistakes, independently of $T$, for some absolute constant $c>0$. Thus, for every fixed class of finite Littlestone dimension when $\delta=\Theta\br{1/\log T}$, private learning requires $\Om{\br{\log T}^{2/3}}$ expected mistakes, whereas private prediction admits $\bigO{\br{\log\log T}^2}$.