🤖 AI Summary
This study addresses the challenging problem of stochastic nonlinear optimization with deterministic equality constraints under heavy-tailed noise characterized by unbounded variance. To this end, we propose TR-SSQP, a method grounded in a trust-region framework and stochastic sequential quadratic programming. By leveraging a normal-tangential decomposition to balance optimality and feasibility, combined with radius normalization and Polyak momentum for stable updates, the algorithm achieves robust performance. The core contribution lies in establishing the first global almost-sure convergence theory without resorting to gradient clipping, thereby filling a critical theoretical gap in this setting. Numerical experiments demonstrate that the proposed approach significantly outperforms existing constrained stochastic optimization algorithms.
📝 Abstract
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrained stochastic methods predominantly rely on bounded-variance assumptions, leaving the heavy-tailed noise regime largely unexplored. To address this gap, we propose a novel trust-region method within the stochastic sequential quadratic programming framework, termed TR-SSQP. Our method employs a normal-tangential decomposition in the step computation to balance optimality and feasibility. In addition, we incorporate a normalization mechanism in the design of the trust-region radius, together with Polyak momentum for gradient estimation, ensuring stable updates without gradient clipping. When the trust-region radius and the momentum parameter decay at appropriate rates, we establish global almost-sure convergence of the method. To the best of our knowledge, this is the first asymptotic convergence result for constrained stochastic optimization under heavy-tailed noise. We demonstrate the promising performance of the proposed method through extensive numerical experiments, including comparisons among its variants and with existing constrained stochastic optimization methods.