🤖 AI Summary
This work addresses the scalability challenge in predict-then-optimize paradigms, where directly minimizing decision regret is hindered by the almost-everywhere non-differentiability of the optimization mapping and reliance on costly solvers. The authors propose a novel training method that eliminates solver calls during learning by designing a surrogate loss grounded in measure transport theory. This approach enables fully solver-free, decision-focused learning while preserving theoretical guarantees—including Fisher consistency and excess risk bounds—and achieves comparable decision performance to state-of-the-art methods. Notably, it reduces training time by several orders of magnitude, marking the first solver-independent framework for predict-then-optimize with both scalability and rigorous theoretical foundations.
📝 Abstract
We propose a scalable method for training prediction (machine learning) models in the predict-then-optimize paradigm, where model outputs serve as coefficients for a subsequent linear optimization task. Directly minimizing the empirical decision regret is intractable for linear programming and combinatorial optimization since the decision mapping is piecewise constant, and the gradients are zero almost everywhere. While existing methods address this by smoothing the differentiation process, they suffer from scalability issues, since a computationally expensive solver call is required for every gradient evaluation. To address this, we propose a decision-focused learning pipeline based on a measure transformation principle, which yields a new surrogate loss that is completely optimization-solver-free during training. We establish theoretical guarantees, including Fisher consistency and excess risk bounds. Empirically, our method achieves decision quality competitive with state-of-the-art methods while reducing training time by orders of magnitude.