Score
Design and implement algorithms that learn resource-allocation policies by formulating allocation as a Markov decision process and applying dynamic programming or MDP-based policy-update methods. This includes building procedures to compute allocation fractions or actions from estimated model parameters and to iteratively alternate estimation and control steps to update the policy.
This paper addresses the challenges of Resource Allocation Optimization (RAO) in dynamic, decentralized environments. To tackle these challenges, it systematically surveys state-of-the-art applications of Multi-Agent Reinforcement Learning (MARL) to RAO. We propose the first three-dimensional taxonomy for RAO—spanning collaboration structure, communication mechanism, and learning paradigm—unifying over 120 recent works and constructing a comprehensive technical landscape across key domains including network slicing, edge computing, and smart grids. By integrating mainstream MARL methodologies—including value decomposition, policy gradient methods, communication-aware learning, and opponent modeling—we establish a method-to-use-case mapping framework. Furthermore, we release an open, continuously updated MARL-RAO research roadmap, accompanied by a technology selection guide and a practical evaluation framework. Our work significantly enhances the deployability of RAO solutions in real-world systems, improving scalability, robustness, and operational feasibility.
This paper addresses a class of multistage stochastic programming problems characterized by continuous state and action spaces, decision-dependent uncertainty, and limited statistical learning capability. To overcome the expressive limitations of conventional models, we propose an extended policy graph framework that explicitly captures the feedback effect of decisions on uncertainty and incorporates online learning mechanisms. Building upon this, we design a novel stochastic dual dynamic programming (SDDP) algorithm and its nonconvex approximation variant, tailored for efficiently solving such structured Markov decision processes. Experimental results on a suite of benchmark instances—increasing in complexity—demonstrate that our approach significantly improves policy quality and computational scalability. The work establishes a new paradigm for stochastic optimization that jointly integrates statistical learning with sequential decision-making, offering both enhanced expressiveness and tractability.
This paper addresses non-cumulative Markov decision processes (NCMDPs), where the objective is to optimize the expectation of an arbitrary function—e.g., maximum reward, Sharpe ratio—of the reward sequence, rather than the conventional discounted cumulative reward. We propose the first general, theoretically rigorous state-augmentation mapping that equivalently transforms any NCMDP into a standard MDP. This reduction enables direct application of classical reinforcement learning algorithms (e.g., DQN, policy gradients) and dynamic programming methods. Empirical evaluation across diverse domains—including control, finance (portfolio optimization), and combinatorial optimization—demonstrates substantial improvements in final performance and training efficiency. Our core contribution is the establishment of a formal theoretical equivalence between NCMDPs and standard MDPs, accompanied by a scalable algorithmic framework for practical implementation. The approach unifies treatment of non-cumulative objectives within the standard RL paradigm while preserving computational tractability and theoretical soundness.
This study addresses the problem of dynamically allocating cores in multicore systems to minimize the steady-state average number of jobs—equivalently, average response time—for two classes of variable-parallelism workloads with unknown speedup parameters. The authors propose an iterative learning-and-control framework that alternates, during job execution, between maximum likelihood estimation of the speedup parameters and updating a core allocation policy derived from a Markov decision process (MDP). Within each class, cores are equally shared among jobs, while the inter-class resource split is determined by the MDP’s optimal solution under the current parameter estimates. This work is the first to integrate online parameter learning with dynamic resource allocation in a closed-loop manner. Numerical experiments demonstrate that the proposed strongly consistent estimator and adaptive scheduling policy significantly reduce the average job count, confirming the estimator’s consistency, policy convergence, and performance gains.
Traditional dynamic memory allocation algorithms—such as first-fit, best-fit, and worst-fit—suffer from excessive fragmentation and poor adaptability under varying request patterns. To address this, this paper proposes the first reinforcement learning (RL)-based adaptive memory management framework. Methodologically, it introduces a history-aware state encoding capturing both free-block distribution and request sequence context, a hierarchical action space, and a customized reward function; it jointly optimizes policy via deep Q-networks (DQN) and policy gradient methods in an end-to-end manner. Key contributions include: (i) the first systematic integration of RL into dynamic memory allocation, and (ii) a history-aware allocation policy that significantly improves generalization under complex and adversarial workloads. Experiments across multiple benchmarks demonstrate substantial improvements over classical algorithms: memory utilization increases by 23% and average fragmentation decreases by 37% under adversarial scenarios.
This paper addresses fair resource allocation in weakly coupled Markov decision processes (MDPs), where $N$ sub-MDPs jointly satisfy a global resource constraint and cannot be optimized independently. Departing from conventional utilitarian objectives—i.e., maximizing total utility—it formalizes fairness via the generalized Gini social welfare function. Theoretically, we establish, for the first time, that under homogeneity, fair optimization is equivalent to maximizing individual utility within the class of permutation-invariant policies. Methodologically, we propose a deep Q-network framework incorporating count-ratio feature encoding, extending fair optimization to heterogeneous settings. Experiments demonstrate that our approach achieves high resource utilization while significantly improving system-level fairness: the Gini coefficient improves by up to 32% compared to baselines.
This study addresses the poor scalability of decision-focused learning in Markov decision processes caused by exhaustive traversal of the entire state space. To overcome this limitation, it proposes an occupancy measure-based linear programming reformulation. The core innovations include introducing an augmented Lagrangian surrogate combined with an occupancy measure LP layer to enable efficient gradient computation, employing randomized row sketching to smooth gradient discontinuities, and designing learnable soft state aggregation alongside neural network function approximation to handle continuous state spaces. Experimental results demonstrate that the proposed method significantly reduces computational costs in multi-task scenarios while achieving lower regret compared to KKT-based baselines and two-stage approaches.
The original TLDR content provided is missing and does not contain specific research information. The following is a standard academic template compliant with the requirements; please replace the bracketed content accordingly: To address the specific challenges and limitations inherent in the core problem, this work proposes a novel method designated as [Method Name]. By leveraging [Core Technical Mechanism 1] and [Core Technical Mechanism 2], the proposed approach effectively resolves the critical bottleneck. Compared to existing baseline models, our method achieves quantitative performance gains on [Evaluation Benchmark], significantly enhancing system characteristics such as robustness and generalization capability. The primary contributions of this study are twofold: it is the first to introduce [Innovation A] into this domain, and it establishes the [Innovation B] framework, thereby providing an efficient and scalable new paradigm for downstream tasks and related research directions.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
This study addresses the problem of fair resource allocation in weakly coupled Markov decision processes with primary and secondary agents, aiming to replace conventional utilitarian objectives with monotonically concave fairness functions. Theoretically, we prove that under symmetry conditions, fairness optimization can be reduced to a specific utilitarian objective. Methodologically, we propose a deep reinforcement learning algorithm based on counting proportions, integrated with a prioritized sampler to achieve efficient solutions. Experimental evaluations on machine replacement and taxi dispatching tasks demonstrate that the proposed approach exhibits both favorable scalability and strong fairness performance.
本文研究了连续时间跳跃马尔可夫决策过程中的强化学习问题,通过建立熵正则化控制问题和开发无模型q学习算法来解决具有通用离散状态空间的应用场景中的探索与利用平衡问题。