Score
Design, build, and evaluate reinforcement learning agents and training pipelines that integrate distributional value estimation into the Soft Actor-Critic family and add a decision assistant module for decision-time guidance; implement the distributional critic, entropy-regularized policy, and the decision-assistant interface and analyze their interactions to jointly optimize trajectory-level objectives and resource constraints, improving sample efficiency, stability, and long-horizon dynamic optimization.
This work addresses risk-sensitive continuous control tasks by proposing a reinforcement learning framework that jointly models the return distribution and policy entropy. Methodologically, it unifies distributional RL and maximum-entropy RL within a single architecture: it explicitly parameterizes the quantile function to model the cumulative reward distribution and extends Soft Actor-Critic with tunable risk measures—such as Conditional Value-at-Risk (CVaR) and entropic risk—to enable flexible control over risk preference (aversion or seeking). The core contribution lies in departing from the conventional expected-return optimization paradigm, enabling end-to-end optimization of arbitrary quantile-based risk metrics. Evaluated on multiple continuous-control benchmarks and risk-sensitive domains—including obstacle avoidance and energy-efficient control—the approach consistently outperforms state-of-the-art methods, achieving superior stability and policy robustness.
This work addresses non-expected utility optimization problems—such as risk-sensitive decision-making and steady-state regulation—where optimizing distributional properties of returns (e.g., tail risk) is critical. We propose a distributed dynamic programming framework incorporating state augmentation, wherein historical reward statistics—specifically, the Conditional Value-at-Risk (CVaR)—are explicitly encoded as augmented state components. Leveraging distributed value and policy iteration, our method directly optimizes statistical functionals of the return distribution without assuming existence or finiteness of expectations. Theoretically, we establish convergence guarantees, derive conditions for objective optimizability, and quantify error bounds for the distributed iterative updates. Empirically, our approach significantly outperforms standard DQN across diverse risk-control and stability benchmarks, achieving superior tail-risk mitigation while preserving steady-state performance.
To address Q-value overestimation—leading to suboptimal policies—in model-free reinforcement learning, this paper proposes DSAC-T: the first algorithm integrating expectation substitution, twin distributional modeling, and variance-aware gradient adjustment within the Soft Actor-Critic (SAC) framework for dual value distribution learning in continuous action spaces. DSAC-T employs variance-weighted gradient clipping and updates to significantly mitigate training instability induced by stochastic returns and sensitivity to reward scaling. Empirical evaluation across multiple standard benchmarks demonstrates that DSAC-T outperforms SAC, TD3, and DDPG without hyperparameter tuning; it exhibits enhanced training stability, strong robustness to reward scaling, and successful deployment on a real-world wheeled robot control task.
In continuous control tasks, modeling only the mean of the state-action value function leads to insufficient policy robustness. This work first observes that the state-action value distribution is highly approximately Gaussian. Leveraging this insight, we propose Normal Quantile Distributional Reinforcement Learning (NQRL): a lightweight variance network estimates the distribution’s standard deviation; Gaussian target quantiles are derived in closed form; and a novel policy update rule is designed to enforce distributional structural consistency. NQRL avoids ensemble-based uncertainty estimation, substantially reducing both parameter count and training overhead. Evaluated on 16 standard continuous control benchmarks, NQRL achieves statistically significant performance improvements on 10 tasks, while converging faster and requiring fewer parameters than state-of-the-art ensemble-based distributional RL methods.
This paper investigates the theoretical advantages of distributional reinforcement learning (DRL) over classical RL, focusing on its implicit environmental exploration capability. Method: We provide the first rigorous decomposition of the distributional loss in categorical DRL, revealing an intrinsic, uncertainty-aware entropy regularization mechanism—spontaneously induced by the structure of the return distribution and requiring no explicit design. This adaptive regularizer transforms environmental uncertainty into enhanced reward signals for policy optimization. Unlike maximum-entropy RL, which explicitly encourages action-space diversity, this mechanism enables implicit, environment-driven exploration grounded in distributional shape. Contribution/Results: Our theoretical analysis uncovers the fundamental reason behind DRL’s superiority over classical RL. Empirical evaluation demonstrates that this implicit regularization significantly improves sample efficiency and policy robustness across diverse benchmarks.
To address policy instability and poor generalization in large language model (LLM) post-training caused by noisy or incomplete reinforcement learning (RL) supervision, this paper proposes a distributed risk-aware RL framework. Methodologically, it introduces Conditional Value-at-Risk (CVaR) theory into token-level distributional value modeling for the first time, and designs an asymmetric risk regularization: contracting the lower tail to suppress noise-induced deviations while preserving the upper tail to retain exploratory diversity. This balances robustness against over-conservatism, thereby enhancing policy generalization. Experiments across multi-turn dialogue, mathematical reasoning, and scientific question answering demonstrate that our method consistently outperforms PPO, GRPO, and robust Bellman-PPO under noisy supervision—achieving superior stability and cross-task transferability.
Traditional reinforcement learning commonly employs diagonal Gaussian policies, which struggle to capture multimodal optimal behaviors and optimize only the mean of the return distribution, thereby neglecting its full structural information and limiting policy performance. This work introduces flow matching into policy modeling for the first time, integrating it with distributional reinforcement learning to construct a policy representation capable of accurately fitting complex, multimodal return distributions. By directly optimizing the entire return distribution to guide policy updates, the proposed method achieves significant performance gains over existing algorithms on MuJoCo continuous control benchmarks, demonstrating not only state-of-the-art results but also enhanced expressiveness in representing policy-induced return distributions.
This study addresses the challenges of deploying Actor-Critic algorithms in real-world control systems, where poor reliability and high sensitivity to hyperparameters often hinder practical application. Focusing on a real-world water treatment plant control task, the authors conduct over 33,000 large-scale ablation experiments to systematically evaluate how key algorithmic components—such as policy update schemes, action distribution representations, gradient estimation methods, and update frequencies—affect performance stability and hyperparameter robustness. Their empirical analysis reveals, for the first time, that commonly adopted default configurations (e.g., Gaussian action distributions with pathwise derivatives) exhibit low reliability, whereas bounded action distributions combined with adaptive update strategies substantially enhance robustness. The work identifies high-stability algorithmic configurations that significantly reduce performance variance under limited tuning budgets, offering actionable, component-level design guidelines for industrial deployment.
This work addresses the challenges of low sample efficiency and training instability in reinforcement learning under stochastic or noisy environments by proposing Distributional Sobolev Training. It introduces a novel approach that jointly models the state-action value function and its gradient as a distribution, constructing a distributional Bellman operator grounded in a first-order world model. The paper establishes the existence of a unique fixed point for this operator, revealing a smoothness trade-off inherent in gradient-aware reinforcement learning. The method employs a conditional variational autoencoder (cVAE) to model the environment dynamics and reward distributions, combined with Max-sliced Maximum Mean Discrepancy (MMD) for distributional Bellman updates. Empirical evaluations on stochastic toy tasks and multiple MuJoCo benchmarks demonstrate significant improvements over existing methods such as MAGE, confirming the approach’s effectiveness and robustness.
This work addresses a critical limitation in conventional federated reinforcement learning, which typically aggregates policies or value functions via parameter averaging and thereby overlooks the multimodality and tail characteristics of reward distributions—leading to performance degradation in safety-critical scenarios. To overcome this, we propose FedDistRL, the first federated distributional reinforcement learning framework, which federates only the quantile-based distributional critic. We further introduce TR-FedDistRL, a novel method that constructs a distributional trust region around local Wasserstein barycenters using a shrink-squash operation, effectively preserving essential statistical properties of the return distribution. Experiments demonstrate that our approach substantially mitigates the mean-blurring effect, reduces safety risks such as accident rates, and alleviates both critic and policy drift, outperforming existing mean-focused and non-federated baselines.