Smooth Learning with Hard Constraints via Legendre-Regularized Policies

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in contextual optimization of simultaneously achieving expressive policy representation, strict satisfaction of hard constraints, and sufficient smoothness for gradient-based training. The authors propose a Legendre-regularized policy that models decisions as solutions to a regularized optimization problem over the original feasible set. This formulation yields a class of policies that inherently satisfy hard constraints, are infinitely differentiable with respect to implicit parameters, and possess universal approximation capability. The approach unifies explicit regularization and implicit perturbation-based smoothing within a single framework. Theoretical analysis leverages implicit function differentiation, Lipschitz continuity, and approximation theory. Empirical results on contextual newsvendor and resource allocation tasks demonstrate significant improvements over existing baselines.
📝 Abstract
We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty terms, and should remain smooth enough for gradient-based training on downstream decision losses. Existing approaches usually emphasize only part of these requirements. We propose Legendre-regularized policies, which parameterize decisions as solutions of regularized optimization problems over the original feasible region. This construction yields policies that are feasible by construction and differentiable with respect to learned latent parameters. We prove that the associated optimizer map is single-valued, maps onto the relative interior of the feasible set, admits an explicit Jacobian, is Lipschitz continuous, and can be made arbitrarily smooth. We also establish a universal approximation result showing that the proposed class can approximate any continuous feasible policy on compact context sets. The framework unifies explicitly regularized optimizers and implicit perturbation-based smooth optimizers. Experiments on contextual newsvendor and resource allocation problems show that our approach improves prescriptive performance relative to the benchmark methods.
Problem

Research questions and friction points this paper is trying to address.

contextual optimization
hard constraints
policy class design
smoothness
feasibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Legendre regularization
hard constraints
differentiable optimization
policy approximation
smooth learning
🔎 Similar Papers