Constrained Learning with Universally Learnable Concept Classes

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the learnability of constrained statistical learning in infinite-dimensional non-convex hypothesis classes. By constructing a dual algorithm within a general reproducing kernel Hilbert space (RKHS) and integrating a norm-ball growth strategy with Tikhonov regularization, the approach ensures strong Lagrangian duality while controlling generalization error. The work introduces two novel concepts—Tikhonov complexity and the closure-realizability gap—and, for the first time, establishes a theory of exact and approximate Probably Approximately Constraint-Consistent (PACC) learnability in non-convex infinite-dimensional settings. Under a source condition, the required sample complexity to achieve ε-optimality scales polynomially in 1/ε; furthermore, when the closure-realizability gap vanishes, exact constraint satisfaction becomes attainable.
📝 Abstract
We study constrained statistical learning over infinite-dimensional hypothesis classes in the fully nonconvex setting, and establish universal PACC learnability of the solutions of dual algorithms: Probably Approximately Correct on Constraints, guaranteeing optimality and constraint satisfaction at once. This strengthens near-PACC results, whose feasibility residual no amount of data can remove. Optimality is caught between generalization, governed by Rademacher complexity and favoring small classes, and strong Lagrangian duality, which rests on Lyapunov convexity for vector measures and needs decomposability, a demand pulling the other way. We reconcile the two by posing the population problem over a universal RKHS $\mathcal{H}_K$, dense in a decomposable envelope, and learning over norm balls of growing radius. This yields the Tikhonov complexity $\mathfrak{T}^{\varepsilon}_{n}$, the least RKHS norm reaching an $\varepsilon$-optimal Lagrangian level set; we prove it finite, obtain exact learnability of the optimal value, and make the sample threshold explicit and polynomial in $1/\varepsilon$ under a source condition. Feasibility is harder: absent convexity the Lagrangian may not attain its infimum, and dual information pins down only an averaged constraint-risk vector, not the risks of any returned predictor. We introduce the closure-realization gap $\varepsilon^\star_\infty$, an index of how well $\mathcal{H}_K$ retrieves feasible solutions from dualization; it is a property of the problem, not of a modeling choice. Learnability is exact when $\varepsilon^\star_\infty=0$, in particular under dual differentiability, and near-PACC with residual exactly $\varepsilon^\star_\infty$ otherwise. Finally, no distribution-free threshold exists already in the unconstrained specialization, so universality is the canonical frame for dual algorithms over large hypothesis classes.
Problem

Research questions and friction points this paper is trying to address.

constrained learning
nonconvex optimization
universal learnability
dual algorithms
feasibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

constrained learning
universal RKHS
PACC learnability
Tikhonov complexity
closure-realization gap