🤖 AI Summary
This study addresses the lack of theoretical convergence comparisons between sketch reuse and fresh sampling in large-scale iterative ridge regression by proposing a residual-aware sampling framework. Through error analysis along the current residual direction, we demonstrate that fresh sketches yield tighter convergence guarantees than those derived from uniform analyses. Furthermore, we derive an oracle distribution based on variance minimization alongside its mixture approximation strategy, integrating leverage score sampling with Qwen2.5 representation probes. Experiments on both synthetic and real-world datasets show that the proposed method significantly accelerates convergence, effectively validating its theoretical and practical advantages.
📝 Abstract
Over the past 25 years, sketching and sampling have become widely used tools for accelerating large-scale regression. In iterative randomized solvers, a basic design choice is whether to $\textit{reuse}$ the same sketch or draw $\textit{fresh}$ randomness at every step. For (under-constrained) iterative ridge regression with column sampling, whether fresh sketches offer provable advantages has remained open: $\textit{We show that they do.}$ Fresh sketching lets us analyze error only along the current residual solution, rather than uniformly over the entire Gram matrix. This directional view yields sharper convergence guarantees for leverage score and ridge leverage score sampling and, more importantly, leads to residual-aware sampling rules. By minimizing the variance of the relevant sketched matrix-vector product, we derive an oracle distribution and practical approximations to the oracle distribution, including a mixture sampling distribution with (somewhat weaker) convergence guarantees. Experiments on synthetic and real data, including ridge probes on Qwen2.5 representations, support our theory, showing substantially faster convergence.