Score
Uses the Karush–Kuhn–Tucker (KKT) optimality conditions to formulate, analyze, and verify solutions of constrained nonlinear optimization problems. Tasks include constructing KKT systems, deriving stationarity, primal and dual feasibility and complementarity conditions, checking necessary and (when applicable) sufficient optimality, performing sensitivity/dual analysis, and using these conditions to guide algorithm or solver design.
This work addresses the ill-conditioning and numerical instability inherent in traditional primal-dual interior-point methods for quadratic programming, which arise from explicitly enforcing complementarity conditions. To overcome this limitation, the authors propose a novel approach that implicitly satisfies the Karush–Kuhn–Tucker (KKT) complementarity conditions. By introducing auxiliary variables, employing a retraction mapping, and replacing the exponential map with the softplus function, the method ensures spectral boundedness of the KKT system, thereby fundamentally mitigating severe ill-conditioning near the solution. Coupled with a linear solver strategy that avoids matrix refactorization at each iteration, the proposed framework not only supports high-accuracy solutions and low-precision arithmetic but also opens new avenues for decomposition-free or indirect solution techniques tailored to large-scale quadratic programming problems.
This paper investigates the computational complexity of computing Karush–Kuhn–Tucker (KKT) points for nonconvex quadratic programming over the unit hypercube $[0,1]^n$. Although KKT conditions constitute only first-order necessary optimality conditions—and are traditionally regarded as computationally easier than global optimization—we establish, for the first time, that computing KKT points is intrinsically hard: the problem is complete for the complexity class CLS (Continuous Local Search). Via a carefully constructed polynomial-time reduction from a CLS-complete problem to KKT point computation, and by integrating analytical tools from both PPAD and PLS, we rigorously prove that KKT point computation is computationally equivalent in difficulty to global optimization. This result breaks the conventional complexity paradigm focused solely on global optima, providing the first tight computational lower bound for first-order critical point computation and revealing the inherent computational robustness—i.e., intractability—of necessary optimality conditions in nonconvex optimization.
For real-time parametric optimization problems (e.g., model predictive control), this paper proposes an end-to-end self-supervised neural iterative solver: a neural network first generates high-quality initial points, which are then refined by a differentiable primal-dual iterative module. The key contributions are twofold: (i) the design of the first KKT-based, label-free loss function, whose global minima are theoretically guaranteed to coincide exactly with KKT points; and (ii) a local convexification approximation strategy for non-convex problems, extending convergence guarantees to non-convex settings. The method requires no ground-truth labels and enables purely self-supervised training. Evaluated on two canonical non-convex benchmark tasks, it achieves a 10× speedup over IPOPT while attaining solution accuracy orders of magnitude higher than existing learning-based approaches.
There exists a significant gap between the theoretical convergence guarantees of deep learning optimization algorithms and their empirical performance, largely due to commonly adopted assumptions—such as Hessian boundedness—that lack empirical validation. Method: We introduce the first trajectory-aware measurement framework tightly aligned with key theoretical quantities, systematically evaluating the validity of mainstream assumptions across diverse architectures and datasets using large-scale training runs. Our framework quantifies dynamic properties—including gradient norms, Hessian spectral characteristics, and loss curvature—along optimization trajectories. Contribution/Results: We find that all examined theoretical assumptions fail to reliably predict actual convergence behavior and exhibit no robust correlation with optimization performance. This work uncovers a fundamental misalignment between theoretical modeling and practice, establishing the first reproducible benchmark for empirically calibrating and reconstructing optimization theory.
This paper addresses constrained stochastic nonlinear optimization problems arising in online statistical inference. We propose Sketch-StoSQP, a sketched stochastic sequential quadratic programming method. Our key contributions are threefold: (i) We establish, for the first time, the asymptotic normality of StoSQP iterates under controllable, non-vanishing approximation errors—ensuring stable per-iteration computational complexity; (ii) We design a plug-and-play covariance estimator enabling immediate statistical inference without algorithmic modification; (iii) We prove that the scaled residual sequence converges in distribution to a non-degenerate zero-mean Gaussian. Empirical evaluation on the CUTEst benchmark and constrained regression tasks demonstrates both statistical validity—accurate coverage rates and well-calibrated confidence intervals—and computational efficiency—constant per-iteration cost and significant overall speedup.
研究了在约束值从样本估计时,增广原始-对偶动力学的稳定性和收敛性问题,通过递归估计约束值来解决均衡偏差。
This work addresses the limitation of existing physics-informed neural networks (PINNs), which enforce nonlinear equality constraints only as soft penalties during inference, thereby failing to guarantee strict feasibility. To overcome this, the authors propose a novel approach that integrates Karush–Kuhn–Tucker (KKT) conditions with a piecewise linear projection mechanism, extending hard-constraint enforcement—previously limited to simpler constraints—to general nonlinear equality settings. By employing an orthogonal projection, the network outputs are rigorously mapped onto the feasible manifold defined by the constraints. Evaluated on a continuous stirred-tank reactor (CSTR) benchmark, the method achieves prediction accuracy comparable to standard PINNs while substantially reducing constraint violations. Moreover, it demonstrates enhanced robustness and lower root-mean-square error (RMSE) under data-scarce conditions.
本文提出了一种基于KKT点拉格朗日乘子特征的非凸优化分类方法,定义了五种操作模式,并通过数值实验验证了理论预测。
This study addresses the challenge of identifying and eliminating redundant contextual directions that exert no practical influence on final decisions in high-dimensional constrained optimization. By leveraging Karush-Kuhn-Tucker conditions and solution sensitivity analysis, this work reveals that active constraints do not necessarily affect optimal decisions. It constructs a decision-preserving interface and introduces a direction-ranking mechanism based on operational domain aggregation to precisely isolate ineffective input dimensions absorbed by dual variables. The proposed approach achieves efficient dimensionality reduction and accurately recovers decision-relevant interfaces under controlled experiments. Notably, it reduces the regret of linear predictors from 0.475 to 0.009, substantially enhancing both decision efficiency and accuracy.
This work addresses a class of nonconvex constrained optimization problems where both the objective and inequality constraints are compositions of convex Lipschitz outer functions with smooth inner mappings. The authors propose a smoothed proximal linear augmented Lagrangian method, reformulating the original problem as a nonsmooth nonconvex-concave minimax problem by restricting dual variables to a compact set. A finite-step mechanism is designed to map stationary points of the truncated minimax problem to KKT points of the original problem. Under a local cone regularity condition, they show that the artificial dual truncation automatically deactivates near feasible points, thereby establishing explicit convergence rates for the KKT residual: a global rate of $O(K^{-1/3})$ under dual regularization, which improves to $O(K^{-1/2})$ when the outer functions are piecewise linear and a local dual error bound holds.