🤖 AI Summary
This study addresses the prohibitive computational cost of decision-focused learning arising from its reliance on expensive optimal reference decisions. To overcome this limitation, we propose a derivative-free training framework based on forward evaluation. By perturbing prediction parameters and directly evaluating the resulting decision costs, the method leverages optimal transport algorithms to determine optimization directions. Furthermore, an infeasibility penalty mechanism is introduced to handle constraints, entirely eliminating the need for both optimal reference decisions and gradient computations throughout the process. We provide rigorous theoretical guarantees for the proposed approach. Extensive experiments on classical optimization benchmarks and real-world tasks demonstrate that our method achieves highly competitive performance without task-specific tuning, significantly lowering the deployment barrier for decision-focused learning.
📝 Abstract
Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predictions through repair or infeasibility penalties. To handle piecewise-constant losses, PolyStepOR perturbs predictor parameters, evaluates the resulting decisions, and uses optimal transport to favor lower-cost directions, requiring no derivatives. Without task-specific tuning, PolyStepOR performs strongly on classical optimization benchmarks and competitively on predicted-constraint and real-world problems. Theoretically, we characterize decision-preserving perturbations and boundary detection, bound sensitivity to cost errors, and establish stationarity guarantees for a smoothed objective. PolyStepOR thus replaces optimal reference decisions and derivatives with forward evaluations.