on-policy decision-focused learning

Design and implement algorithms and training pipelines that learn predictive models or policies from on-policy (sequentially collected) data using decision-focused objectives so the learned component directly maximizes downstream decision quality under partial or bandit-like feedback. Analyze and provide convergence and performance guarantees for these methods (for example policy-gradient rates such as o(t^{-1/2})), and build the data-collection and feedback-integration mechanisms required for on-policy operation.

on-policydecision-focusedlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

From Sequential to Recursive: Enhancing Decision-Focused Learning with Bidirectional Feedback

Nov 11, 2025
XW
Xinyu Wang
🏛️ The Hong Kong Polytechnic University

Existing sequential decision-focused learning (S-DFL) suffers from unidirectional prediction→optimization pipelines, limiting its ability to model bidirectional feedback in complex interactive decision-making scenarios. To address this, we propose recursive decision-focused learning (R-DFL), the first framework to establish a closed-loop, iterative interaction between prediction and optimization modules, enabling end-to-end joint modeling. R-DFL unifies explicit unrolling with implicit differentiation based on fixed-point equations, ensuring both high-precision and computationally efficient gradient propagation through the optimization layer. Empirically, R-DFL achieves significant improvements over S-DFL on benchmark tasks—including the newsvendor problem and bipartite matching—yielding markedly superior decision quality. Moreover, it demonstrates strong cross-scenario generalization capability, validating its robustness beyond task-specific training distributions.

Developing efficient gradient methods for recursive decision-making systemsEnhancing decision quality through bidirectional prediction-optimization feedbackOvercoming sequential limitations in decision-focused learning frameworks

This work systematically investigates decision-focused learning (DFL) in the context of stochastic linear programming, revealing that under the conventional “predict-then-optimize” paradigm, improved prediction accuracy does not necessarily translate into better downstream decision quality. The study demonstrates fundamental limitations of standard statistical learning approaches and common data collection strategies—along with distributional metrics such as Wasserstein distance—when applied to optimization tasks. By developing a unified framework that jointly models prediction and optimization, the paper clarifies the essential differences between DFL and traditional predictive modeling. Building on these insights, it proposes novel methods explicitly designed to optimize decision performance, thereby laying foundational groundwork for theory and tools in decision-oriented machine learning.

decision qualitydecision-focused learningpredict-then-optimize

This work addresses decision-focused learning in sequential contextual linear optimization under partial feedback. The authors propose an online policy learning method that jointly optimizes a predictive model and the downstream linear decision task through a stochastic predict-and-optimize framework. The key innovation lies in extending decision-focused learning to the online, partial-feedback setting for the first time and introducing a hybrid gradient estimator that integrates a scoring function with a plug-in architecture, effectively leveraging structural information from the downstream optimization problem. Empirical results demonstrate that the proposed approach significantly outperforms contextual bandit baselines across multiple tasks—including top-k selection, shortest path, combinatorial pricing, and real-world energy dispatch—achieving substantially lower cumulative regret.

bandit feedbackcontextual linear optimizationdecision-focused learning

Online Decision-Focused Learning

May 19, 2025
AC
Aymeric Capitaine
🏛️ École polytechnique | Inria Paris | Ecole Normale Supérieure | PSL | Inria Saclay | Université Paris Saclay

This work studies decision-focused learning (DFL) in dynamic environments, where the objective function is nonsmooth (with zero or undefined gradients) and nonconvex, while data distributions and time-varying constraints evolve continuously—rendering conventional DFL approaches ineffective. We first extend DFL to an online learning framework and propose a synergistic mechanism combining differentiable regularization with optimistic prediction. Leveraging a near-optimal oracle and dynamic regret analysis, we establish the first theoretical guarantee of bounded expected dynamic regret for time-varying constraints—valid over simplex and convex polyhedral decision spaces. Experiments on time-varying knapsack problems demonstrate that our method significantly outperforms existing prediction-focused baselines, achieving both rigorous theoretical foundations and strong empirical performance.

Addressing dynamic environments in Decision-Focused Learning (DFL)Developing an online algorithm for DFL with dynamic regret boundsHandling non-differentiable and non-convex objective functions in DFL

Score-Aware Policy-Gradient Methods and Performance Guarantees using Local Lyapunov Conditions: Applications to Product-Form Stochastic Networks and Queueing Systems

Dec 05, 2023
CC
Céline Comte
🏛️ CNRS | LAAS | Eindhoven University of Technology | IRIT | Université Toulouse III Paul Sabatier

This paper addresses average-reward reinforcement learning for countable-state Markov decision processes (MDPs) with potentially unstable policies, where the stationary distribution belongs to an exponential family parameterized by the policy. Method: We propose the Score-Aware Gradient Estimator (SAGE), a value-function-free policy gradient estimator that directly exploits the exponential-family structure of the stationary distribution—bypassing the conventional actor-critic reliance on value function approximation. Contribution/Results: Theoretically, under non-convexity and infinite state spaces, we establish convergence guarantees via local Lyapunov conditions and Hessian non-degeneracy. Empirically, on multi-class product-form stochastic networks and queueing systems, SAGE achieves significantly faster training convergence to near-optimal policies compared to standard actor-critic methods, thereby validating both theoretical soundness and practical efficacy.

Ensuring policy convergence using local Lyapunov stability analysisEstimating gradients without value-function approximation in MDPsImproving policy-gradient methods for model-based reinforcement learning

Latest Papers

What's happening recently
View more

This study addresses the lack of theoretical foundations for off-policy self-generated data fine-tuning in large model post-training, where convergence is difficult to guarantee when the sampling distribution is updated infrequently. To this end, this work proposes RE(S), a unified framework that formulates the optimization as a staged KL-divergence minimization process and provides rigorous analysis by integrating multi-armed bandits with a generalized REINFORCE algorithm under a Softmax policy. The authors prove that global convergence is achievable for any fixed S, establishing a tight O(1/T) convergence rate. Furthermore, they reveal a distinct advantage of off-policy learning: under weak initialization, appropriately increasing S helps escape local traps and significantly accelerates convergence, thereby breaking the conventional reliance on strict on-policy training.

convergence ratefine-tuninglearning dynamics

This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.

function approximationMarkov decision processesmathematical foundations

When data are insufficient to learn policies with low regret or significantly better performance than a baseline, how can we characterize the intrinsic difficulty of policy learning? This work proposes a unified framework to systematically study three fundamental problems: optimal policy learning, improved policy learning, and policy existence verification. Through theoretical analysis, problem reductions, and sample complexity comparisons, the paper establishes a strict or partially strict hierarchy of difficulty among these tasks: optimal policy learning is provably harder than improved policy learning, and under natural conditions, a sublinear polynomial complexity gap separates improved policy learning from existence verification. Notably, this study formalizes the policy existence problem for the first time and reveals that even when constructing an improved policy is infeasible, efficiently determining its existence may still be possible.

improving policyoptimal policypolicy existence

Offline reinforcement learning typically relies on stepwise rewards, yet real-world datasets often provide only trajectory-level labels, creating a statistical efficiency bottleneck for policy optimization. This work proposes OPAC, an algorithm that combines implicit reward modeling with a pessimistic Actor-Critic framework to enable effective policy learning under trajectory-level supervision alone, and extends it to settings with preference feedback and generalized trajectory objectives. The study establishes the first statistical theory for this setting, providing matching upper and lower bounds on sample complexity, revealing the fundamental challenge posed by the absence of stepwise rewards, and identifying structural conditions under which efficient learning is achievable. Under both standard and preference-based settings, OPAC attains an error bound of Õ(H²√(C_sa(π*)/n)); when the identified structural conditions hold, the generalized OPAC achieves polynomial sample complexity.

offline reinforcement learningoutcome-based feedbacksample efficiency

This work addresses value-based policy learning in operational settings characterized by limited data and a continuous, high-dimensional state-action space, achieving rapid regret convergence through greedy policies induced by Q*-function estimates. The core contribution lies in identifying three geometric structures governing convergence rates in continuous action spaces: the growth exponent \( p \), the boundary quality exponent \( m \), and the action regularity exponent \( q \). It is rigorously shown that when \( q > 0 \), policy regret converges faster than \( n^{-1/2} \). By integrating minimax analysis with Q* estimation, the theoretical results are validated in practical scenarios such as dynamic inventory management and service allocation, establishing minimax-optimal regret convergence rates and providing a solid theoretical foundation for efficient learning in high-dimensional continuous decision-making problems.

continuous action spacesfast convergencepolicy regret

Hot Scholars

YL

Yining Li

Shanghai AI Laboratory
Multimodal LearningLarge Language Model
JA

Jon Ander Campos

Cohere
Natural Language ProcessingDeep LearningMachine Learning
WY

Wenming Yang

Tsinghua University
Computer VisionImage Processing
WB

Walid Bousselham

University of Bonn
MultimodalMachine learningComputer Vision