When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of flatness proxies in robustness certification and training interventions, noting that while curvature upper bounds are valid, they remain inequivalent to intrinsic model properties. To resolve this, we propose a globally valid, gauge-invariant feature space rectification method that leverages gauge-invariant geometric analysis and softmax symmetry handling. Specifically, we establish row-centering as an orbit-minimizing bound, overcoming the inability of scalar rescaling to align probability updates. Extensive long-training experiments on CIFAR-10 and 45 paired tests demonstrate that the quotient-regularized predictor maintains alignment. Furthermore, this correction significantly and reversibly suppresses generalization degradation, effectively restoring predictive consistency.
📝 Abstract
A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property. We demonstrate this distinction for a last-layer relative-flatness proxy used in both settings. First, empirical-risk stationarity does not eliminate pointwise first-order loss terms: at a finite global empirical-risk minimum, the retained certificate expression underestimates a loss increase by over $210\times$. We derive a globally valid, gauge-invariant feature-space repair. Second, common-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded. Even standard reference-class choices double it on average relative to the centered representation. For a single fixed-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates. Row centering gives the orbit-minimized bound and restores value and full-model gradient invariance under this symmetry. Across 45 paired one-step tests on algorithmic and image models, amplified shifts separate raw-regularized predictors while quotient-regularized predictors remain aligned. Long-horizon CIFAR-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent. Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention.
Problem

Research questions and friction points this paper is trying to address.

robustness certificates
flatness proxy
curvature upper bound
training interventions
softmax shift invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

flatness proxy
robustness certificates
gauge invariance
row centering
quotient regularization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.