🤖 AI Summary
This study addresses the challenges of biased mini-batch gradients and computational constraints in large-scale Cox regression. To overcome these limitations, the authors introduce auxiliary variables and a Softplus surrogate function to eliminate gradient bias, leveraging grouped risk set structures to achieve unbiased single-sample stochastic optimization alongside an improved convergence rate analysis for the LogSumExp approximation. Theoretically, the proposed estimator is proven to attain a mean squared error rate of $T^{-4/5}$, with an asymptotic distribution consistent with that of the full-data estimator. Empirical evaluations further demonstrate that this approach significantly outperforms existing baseline methods.
📝 Abstract
Motivated by the computational challenges of large-scale Cox regression, we study stochastic minimization of LogSumExp objectives over large sets. Mini-batch normalizer estimates generally yield biased gradients. We instead use a softplus surrogate that introduces one auxiliary scalar per normalizer and admits unbiased single-sample gradients. For smooth convex LogSumExp objectives, we prove an $O(T^{-1/2})$ averaged objective bound, improving the previous $T^{-1/4}$ analysis. With a strongly convex regularizer on the original variable, we also obtain a last-iterate squared-error rate of $\widetilde{O}(T^{-1})$ without strong convexity in the auxiliary variables. For Cox regression, the normalizers are defined over nested risk sets. We exploit this structure by grouping neighboring failures and sharing one auxiliary variable per group. The resulting compressed objective admits uniform score and curvature bounds that control the errors from grouping and softplus approximation. Together with the general optimization result, these bounds give a mean-square rate of $T^{-4/5}$, up to logarithmic factors, relative to the full Cox solution. The compressed estimator also matches the full estimator's asymptotic distribution. Experiments on synthetic and real survival datasets with slowly decreasing risk sets show a favorable performance relative to stochastic baselines.