learn allocation policy

Design and implement algorithms that learn resource-allocation policies by formulating allocation as a Markov decision process and applying dynamic programming or MDP-based policy-update methods. This includes building procedures to compute allocation fractions or actions from estimated model parameters and to iteratively alternate estimation and control steps to update the policy.

learnallocationpolicy

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

MDP modeling for multi-stage stochastic programs

Sep 26, 2025
DP
David P. Morton
🏛️ Northwestern University | Dowson Farms | SKEMA Business School | Université Côte d’Azur

This paper addresses a class of multistage stochastic programming problems characterized by continuous state and action spaces, decision-dependent uncertainty, and limited statistical learning capability. To overcome the expressive limitations of conventional models, we propose an extended policy graph framework that explicitly captures the feedback effect of decisions on uncertainty and incorporates online learning mechanisms. Building upon this, we design a novel stochastic dual dynamic programming (SDDP) algorithm and its nonconvex approximation variant, tailored for efficiently solving such structured Markov decision processes. Experimental results on a suite of benchmark instances—increasing in complexity—demonstrate that our approach significantly improves policy quality and computational scalability. The work establishes a new paradigm for stochastic optimization that jointly integrates statistical learning with sequential decision-making, offering both enhanced expressiveness and tractability.

Develops new stochastic dual dynamic programming variantsExtends MDP modeling for multi-stage stochastic programsIncorporates decision-dependent uncertainty in transition probabilities

Tackling Decision Processes with Non-Cumulative Objectives using Reinforcement Learning

May 22, 2024
MN
Maximilian Nägele
🏛️ Max Planck Institute for the Science of Light | Friedrich-Alexander-Universität Erlangen-Nürnberg

This paper addresses non-cumulative Markov decision processes (NCMDPs), where the objective is to optimize the expectation of an arbitrary function—e.g., maximum reward, Sharpe ratio—of the reward sequence, rather than the conventional discounted cumulative reward. We propose the first general, theoretically rigorous state-augmentation mapping that equivalently transforms any NCMDP into a standard MDP. This reduction enables direct application of classical reinforcement learning algorithms (e.g., DQN, policy gradients) and dynamic programming methods. Empirical evaluation across diverse domains—including control, finance (portfolio optimization), and combinatorial optimization—demonstrates substantial improvements in final performance and training efficiency. Our core contribution is the establishment of a formal theoretical equivalence between NCMDPs and standard MDPs, accompanied by a scalable algorithmic framework for practical implementation. The approach unifies treatment of non-cumulative objectives within the standard RL paradigm while preserving computational tractability and theoretical soundness.

Enabling RL techniques to optimize arbitrary reward functions in NCMDPsImproving performance and training efficiency in diverse NCMDP applicationsMapping non-cumulative MDPs to standard MDPs for broader applicability

This study addresses the problem of dynamically allocating cores in multicore systems to minimize the steady-state average number of jobs—equivalently, average response time—for two classes of variable-parallelism workloads with unknown speedup parameters. The authors propose an iterative learning-and-control framework that alternates, during job execution, between maximum likelihood estimation of the speedup parameters and updating a core allocation policy derived from a Markov decision process (MDP). Within each class, cores are equally shared among jobs, while the inter-class resource split is determined by the MDP’s optimal solution under the current parameter estimates. This work is the first to integrate online parameter learning with dynamic resource allocation in a closed-loop manner. Numerical experiments demonstrate that the proposed strongly consistent estimator and adaptive scheduling policy significantly reduce the average job count, confirming the estimator’s consistency, policy convergence, and performance gains.

dynamic core allocationmalleable jobsresource allocation

Reinforcement Learning for Dynamic Memory Allocation

Oct 20, 2024
AL
Arisrei Lim
🏛️ University of Texas at Austin

Traditional dynamic memory allocation algorithms—such as first-fit, best-fit, and worst-fit—suffer from excessive fragmentation and poor adaptability under varying request patterns. To address this, this paper proposes the first reinforcement learning (RL)-based adaptive memory management framework. Methodologically, it introduces a history-aware state encoding capturing both free-block distribution and request sequence context, a hierarchical action space, and a customized reward function; it jointly optimizes policy via deep Q-networks (DQN) and policy gradient methods in an end-to-end manner. Key contributions include: (i) the first systematic integration of RL into dynamic memory allocation, and (ii) a history-aware allocation policy that significantly improves generalization under complex and adversarial workloads. Experiments across multiple benchmarks demonstrate substantial improvements over classical algorithms: memory utilization increases by 23% and average fragmentation decreases by 37% under adversarial scenarios.

Creating adaptive strategies for adversarial memory request patternsDeveloping RL framework for dynamic memory allocation managementOvercoming fragmentation and inefficiency of traditional allocation algorithms

Fair Resource Allocation in Weakly Coupled Markov Decision Processes

Nov 14, 2024
XT
Xiaohui Tu
🏛️ HEC Montréal | MILA - Quebec AI Institute

This paper addresses fair resource allocation in weakly coupled Markov decision processes (MDPs), where $N$ sub-MDPs jointly satisfy a global resource constraint and cannot be optimized independently. Departing from conventional utilitarian objectives—i.e., maximizing total utility—it formalizes fairness via the generalized Gini social welfare function. Theoretically, we establish, for the first time, that under homogeneity, fair optimization is equivalent to maximizing individual utility within the class of permutation-invariant policies. Methodologically, we propose a deep Q-network framework incorporating count-ratio feature encoding, extending fair optimization to heterogeneous settings. Experiments demonstrate that our approach achieves high resource utilization while significantly improving system-level fairness: the Gini coefficient improves by up to 32% compared to baselines.

Fair resource allocation in weakly coupled MDPsGeneralized Gini function for fairness definitionReduction to utilitarian objective in homogeneous cases

Latest Papers

What's happening recently
View more

This study addresses the poor scalability of decision-focused learning in Markov decision processes caused by exhaustive traversal of the entire state space. To overcome this limitation, it proposes an occupancy measure-based linear programming reformulation. The core innovations include introducing an augmented Lagrangian surrogate combined with an occupancy measure LP layer to enable efficient gradient computation, employing randomized row sketching to smooth gradient discontinuities, and designing learnable soft state aggregation alongside neural network function approximation to handle continuous state spaces. Experimental results demonstrate that the proposed method significantly reduces computational costs in multi-task scenarios while achieving lower regret compared to KKT-based baselines and two-stage approaches.

Decision-Focused LearningLinear ProgrammingMarkov Decision Process

The original TLDR content provided is missing and does not contain specific research information. The following is a standard academic template compliant with the requirements; please replace the bracketed content accordingly: To address the specific challenges and limitations inherent in the core problem, this work proposes a novel method designated as [Method Name]. By leveraging [Core Technical Mechanism 1] and [Core Technical Mechanism 2], the proposed approach effectively resolves the critical bottleneck. Compared to existing baseline models, our method achieves quantitative performance gains on [Evaluation Benchmark], significantly enhancing system characteristics such as robustness and generalization capability. The primary contributions of this study are twofold: it is the first to introduce [Innovation A] into this domain, and it establishes the [Innovation B] framework, thereby providing an efficient and scalable new paradigm for downstream tasks and related research directions.

Endogenous Markov StateFinite-horizonLP Re-solving

This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.

function approximationMarkov decision processesmathematical foundations

This study addresses the problem of fair resource allocation in weakly coupled Markov decision processes with primary and secondary agents, aiming to replace conventional utilitarian objectives with monotonically concave fairness functions. Theoretically, we prove that under symmetry conditions, fairness optimization can be reduced to a specific utilitarian objective. Methodologically, we propose a deep reinforcement learning algorithm based on counting proportions, integrated with a prioritized sampler to achieve efficient solutions. Experimental evaluations on machine replacement and taxi dispatching tasks demonstrate that the proposed approach exhibits both favorable scalability and strong fairness performance.

Fair resource allocationFairness optimizationSequential decision-making

Hot Scholars

TZ

Tong Zhao

Zhejiang University & Westlake University
Machine LearningDiffusion Model
DI

Dong In Kim

Sungkyunkwan University (SKKU)
Wireless CommunicationsInternet of ThingsWireless Power TransferConnected Intelligence
YD

Yuwei Du

Tsinghua University
trajectory modelling
WY

Weijie Yuan

SUSTech
OTFSISACLow-Altitude Wireless Network
BZ

Beier Zhu

Research Scientist, Nanyang Technological University
Robust Machine Learning