🤖 AI Summary
This study addresses the oscillatory behavior and slow convergence of gradient descent (GD) with large learning rates in logistic regression over linearly separable data. To characterize the transient dynamics of large-step-size GD for arbitrarily dimensional linearly separable datasets, this work proposes a recursive nested interval decomposition framework. Overcoming the limitations of existing two-dimensional analyses, it bounds the nesting depth using the classification margin and data rank, complemented by combinatorial counting arguments for rigorous theoretical analysis. The main contribution establishes that GD requires only O(ln^p(1/ε)) iterations to reduce the loss to ε. This result significantly improves upon existing convergence rate bounds, offering stronger theoretical guarantees and enhanced efficiency for large-learning-rate training paradigms.
📝 Abstract
We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrt{\epsilon})$ to reach loss $\epsilon$ with an aggressive stepsize, although the loss may initially oscillate. Tighter control of the oscillatory dynamics has been available only for two-dimensional data. We prove a substantially faster rate in arbitrary dimension: GD with a large stepsize $\eta=1/\epsilon$ reaches loss $\epsilon$ within $O(\ln^{p}(1/\epsilon))$ steps, where $p$ depends only on the margin and the rank of the data. Our proof improves the bound on the transition time of GD from the oscillatory to the stable phase, after which the loss decreases monotonically. We split the oscillatory phase into recursively nested intervals. The margin and the rank bound the nesting depth, and a counting argument bounds the number of intervals at each depth, together yielding the polylogarithmic step complexity.