Active Learning For Contextual Linear Optimization: A Margin-Based Approach

📅 2023-05-11
📈 Citations: 6
✨ Influential: 1
📄 PDF
🤖 AI Summary
This work addresses contextual linear optimization, where label acquisition—i.e., querying target function coefficients—is costly, and aims to minimize decision error measured by the Smart Predict-then-Optimize (SPO) loss via active learning. We propose the first SPO-loss-driven active learning framework, which directly incorporates the SPO loss into the query strategy and introduces a novel margin-based selection criterion grounded in distance-to-decision-boundary analysis to dynamically identify the most informative unlabeled instances. Theoretically, we establish the first label complexity upper bound and generalization risk guarantee for SPO and its surrogate SPO+ loss under margin conditions. Empirically, our method substantially reduces labeling effort while outperforming fully supervised baselines on personalized pricing and shortest-path prediction tasks.
📝 Abstract
We develop the first active learning method for contextual linear optimization. Specifically, we introduce a label acquisition algorithm that sequentially decides whether to request the ``labels'' of feature samples from an unlabeled data stream, where the labels correspond to the coefficients of the objective in the linear optimization. Our method is the first to be directly informed by the decision loss induced by the predicted coefficients, referred to as the Smart Predict-then-Optimize (SPO) loss. Motivated by the structure of the SPO loss, our algorithm adopts a margin-based criterion utilizing the concept of distance to degeneracy. In particular, we design an efficient active learning algorithm with theoretical excess risk (i.e., generalization) guarantees. We derive upper bounds on the label complexity, defined as the number of samples whose labels are acquired to achieve a desired small level of SPO risk. These bounds show that our algorithm has a much smaller label complexity than the naive supervised learning approach that labels all samples, particularly when the SPO loss is minimized directly on the collected data. To address the discontinuity and nonconvexity of the SPO loss, we derive label complexity bounds under tractable surrogate loss functions. Under natural margin conditions, these bounds also outperform naive supervised learning. Using the SPO+ loss, a specialized surrogate of the SPO loss, we establish even tighter bounds under separability conditions. Finally, we present numerical evidence showing the practical value of our algorithms in settings such as personalized pricing and the shortest path problem.
Problem

Research questions and friction points this paper is trying to address.

Active Learning
Linear Optimization
Decision Loss Minimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Learning
Smart Predict-then-Optimize (SPO) Loss
Contextual Linear Optimization
💼 Related Jobs
No related jobs found.
University of California, Berkeley
M
Mo Liu
Department of Industrial Engineering and Operations Research, University of California, Berkeley, Berkeley, CA, 94720
P
Paul Grigas
Department of Industrial Engineering and Operations Research, University of California, Berkeley, Berkeley, CA, 94720
H
Heyuan Liu
Department of Industrial Engineering and Operations Research, University of California, Berkeley, Berkeley, CA, 94720
Z
Z. Shen
Department of Industrial Engineering and Operations Research, University of California, Berkeley, Berkeley, CA, 94720