Score
Design and implement algorithms and training pipelines that learn predictive models or policies from on-policy (sequentially collected) data using decision-focused objectives so the learned component directly maximizes downstream decision quality under partial or bandit-like feedback. Analyze and provide convergence and performance guarantees for these methods (for example policy-gradient rates such as o(t^{-1/2})), and build the data-collection and feedback-integration mechanisms required for on-policy operation.
Existing sequential decision-focused learning (S-DFL) suffers from unidirectional prediction→optimization pipelines, limiting its ability to model bidirectional feedback in complex interactive decision-making scenarios. To address this, we propose recursive decision-focused learning (R-DFL), the first framework to establish a closed-loop, iterative interaction between prediction and optimization modules, enabling end-to-end joint modeling. R-DFL unifies explicit unrolling with implicit differentiation based on fixed-point equations, ensuring both high-precision and computationally efficient gradient propagation through the optimization layer. Empirically, R-DFL achieves significant improvements over S-DFL on benchmark tasks—including the newsvendor problem and bipartite matching—yielding markedly superior decision quality. Moreover, it demonstrates strong cross-scenario generalization capability, validating its robustness beyond task-specific training distributions.
This work systematically investigates decision-focused learning (DFL) in the context of stochastic linear programming, revealing that under the conventional “predict-then-optimize” paradigm, improved prediction accuracy does not necessarily translate into better downstream decision quality. The study demonstrates fundamental limitations of standard statistical learning approaches and common data collection strategies—along with distributional metrics such as Wasserstein distance—when applied to optimization tasks. By developing a unified framework that jointly models prediction and optimization, the paper clarifies the essential differences between DFL and traditional predictive modeling. Building on these insights, it proposes novel methods explicitly designed to optimize decision performance, thereby laying foundational groundwork for theory and tools in decision-oriented machine learning.
This work addresses decision-focused learning in sequential contextual linear optimization under partial feedback. The authors propose an online policy learning method that jointly optimizes a predictive model and the downstream linear decision task through a stochastic predict-and-optimize framework. The key innovation lies in extending decision-focused learning to the online, partial-feedback setting for the first time and introducing a hybrid gradient estimator that integrates a scoring function with a plug-in architecture, effectively leveraging structural information from the downstream optimization problem. Empirical results demonstrate that the proposed approach significantly outperforms contextual bandit baselines across multiple tasks—including top-k selection, shortest path, combinatorial pricing, and real-world energy dispatch—achieving substantially lower cumulative regret.
This work studies decision-focused learning (DFL) in dynamic environments, where the objective function is nonsmooth (with zero or undefined gradients) and nonconvex, while data distributions and time-varying constraints evolve continuously—rendering conventional DFL approaches ineffective. We first extend DFL to an online learning framework and propose a synergistic mechanism combining differentiable regularization with optimistic prediction. Leveraging a near-optimal oracle and dynamic regret analysis, we establish the first theoretical guarantee of bounded expected dynamic regret for time-varying constraints—valid over simplex and convex polyhedral decision spaces. Experiments on time-varying knapsack problems demonstrate that our method significantly outperforms existing prediction-focused baselines, achieving both rigorous theoretical foundations and strong empirical performance.
This paper addresses average-reward reinforcement learning for countable-state Markov decision processes (MDPs) with potentially unstable policies, where the stationary distribution belongs to an exponential family parameterized by the policy. Method: We propose the Score-Aware Gradient Estimator (SAGE), a value-function-free policy gradient estimator that directly exploits the exponential-family structure of the stationary distribution—bypassing the conventional actor-critic reliance on value function approximation. Contribution/Results: Theoretically, under non-convexity and infinite state spaces, we establish convergence guarantees via local Lyapunov conditions and Hessian non-degeneracy. Empirically, on multi-class product-form stochastic networks and queueing systems, SAGE achieves significantly faster training convergence to near-optimal policies compared to standard actor-critic methods, thereby validating both theoretical soundness and practical efficacy.
This study addresses the lack of theoretical foundations for off-policy self-generated data fine-tuning in large model post-training, where convergence is difficult to guarantee when the sampling distribution is updated infrequently. To this end, this work proposes RE(S), a unified framework that formulates the optimization as a staged KL-divergence minimization process and provides rigorous analysis by integrating multi-armed bandits with a generalized REINFORCE algorithm under a Softmax policy. The authors prove that global convergence is achievable for any fixed S, establishing a tight O(1/T) convergence rate. Furthermore, they reveal a distinct advantage of off-policy learning: under weak initialization, appropriately increasing S helps escape local traps and significantly accelerates convergence, thereby breaking the conventional reliance on strict on-policy training.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
When data are insufficient to learn policies with low regret or significantly better performance than a baseline, how can we characterize the intrinsic difficulty of policy learning? This work proposes a unified framework to systematically study three fundamental problems: optimal policy learning, improved policy learning, and policy existence verification. Through theoretical analysis, problem reductions, and sample complexity comparisons, the paper establishes a strict or partially strict hierarchy of difficulty among these tasks: optimal policy learning is provably harder than improved policy learning, and under natural conditions, a sublinear polynomial complexity gap separates improved policy learning from existence verification. Notably, this study formalizes the policy existence problem for the first time and reveals that even when constructing an improved policy is infeasible, efficiently determining its existence may still be possible.
Offline reinforcement learning typically relies on stepwise rewards, yet real-world datasets often provide only trajectory-level labels, creating a statistical efficiency bottleneck for policy optimization. This work proposes OPAC, an algorithm that combines implicit reward modeling with a pessimistic Actor-Critic framework to enable effective policy learning under trajectory-level supervision alone, and extends it to settings with preference feedback and generalized trajectory objectives. The study establishes the first statistical theory for this setting, providing matching upper and lower bounds on sample complexity, revealing the fundamental challenge posed by the absence of stepwise rewards, and identifying structural conditions under which efficient learning is achievable. Under both standard and preference-based settings, OPAC attains an error bound of Õ(H²√(C_sa(π*)/n)); when the identified structural conditions hold, the generalized OPAC achieves polynomial sample complexity.
This work addresses value-based policy learning in operational settings characterized by limited data and a continuous, high-dimensional state-action space, achieving rapid regret convergence through greedy policies induced by Q*-function estimates. The core contribution lies in identifying three geometric structures governing convergence rates in continuous action spaces: the growth exponent \( p \), the boundary quality exponent \( m \), and the action regularity exponent \( q \). It is rigorously shown that when \( q > 0 \), policy regret converges faster than \( n^{-1/2} \). By integrating minimax analysis with Q* estimation, the theoretical results are validated in practical scenarios such as dynamic inventory management and service allocation, establishing minimax-optimal regret convergence rates and providing a solid theoretical foundation for efficient learning in high-dimensional continuous decision-making problems.