Score
Designs and trains defender policies using multi-objective reinforcement learning to optimize tradeoffs among competing objectives (for example safety versus utility or false positives versus false negatives). Builds reward and constraint formulations, learning algorithms, and evaluation procedures to analyze Pareto fronts, tune trade-offs, and produce robust defender behavior under differing threat and cost preferences.
Multi-objective reinforcement learning (MORL) suffers from non-unique mappings between policy parameter space and multi-objective performance space, poor interpretability, and low efficiency in Pareto frontier search. To address these challenges, this paper proposes an interpretable MORL framework based on local linear mapping—first embedding bidirectional parameter–performance interpretability into algorithm design. Specifically, it models the parameter-to-performance mapping locally as linear, enabling real-time semantic interpretation of policy objectives; supports zero-shot cross-domain policy transfer without retraining; and facilitates interpretable gradient-guided approximation of the Pareto frontier. Evaluated on multiple benchmark tasks, our method achieves significant improvements: +18.7% in Pareto frontier coverage and 2.3× acceleration in convergence speed, while outperforming state-of-the-art methods in both explanation quality and search efficiency.
This work addresses a critical limitation in multi-objective reinforcement learning, where policy evaluation based solely on value vectors often overlooks behavioral differences, leading to ambiguous decision-making. To resolve this, the authors propose an exploratory diagnostic framework that explicitly incorporates behavioral divergence into Pareto front analysis for the first time. By integrating trajectory clustering with visualization techniques, the method quantitatively reveals the behavioral diversity among Pareto-optimal policies. Empirical validation on both grid-world and continuous control benchmarks demonstrates its effectiveness: even in complex tasks, the framework clearly delineates behavioral distinctions between policies, thereby offering decision-makers a richer, more informative basis for policy selection.
Autonomous defense in military and critical infrastructure networks faces inherent trade-offs among conflicting objectives—such as attack suppression and business continuity assurance—posing significant challenges for conventional single-objective reinforcement learning (RL). Method: This work pioneers the integration of multi-objective reinforcement learning (MORL) into cyber offense-defense games, proposing two online-adaptive autonomous cyber defense (ACD) agents: Multi-Objective Proximal Policy Optimization (MOPPO) and Pareto-Conditioned Network (PCN). Both incorporate Pareto-optimality constraints into the PPO framework and dynamically adapt to operational constraints within a high-fidelity autonomous cyber defense simulation environment. Contribution/Results: Experimental evaluation demonstrates that the proposed MORL agents achieve Pareto-superior balance across attack interception rate, service availability, and recovery cost—outperforming single-objective RL baselines significantly. This advances beyond the fundamental limitation of traditional RL in simultaneously optimizing multiple, often competing, security operation objectives.
This work addresses the challenge in multi-objective reinforcement learning (MORL) where single-policy approaches often fail to fully recover the Pareto front due to gradient interference and policy representation collapse. To overcome this, the authors propose the D³PO framework, which decouples the optimization of individual objectives to preserve distinct learning signals, delays preference fusion, and introduces a scaled diversity regularizer to enhance policy sensitivity to preferences. As the first method to systematically identify and mitigate gradient interference and representation collapse in preference-conditioned policies, D³PO achieves state-of-the-art or comparable performance using only a single deployable policy across multiple MORL benchmarks, significantly improving both the coverage and quality of the recovered Pareto front as measured by hypervolume and expected utility metrics.
This work addresses the challenge in multi-objective reinforcement learning that traditional scalarization methods often fail to fully capture the Pareto optimal front. To overcome this limitation, the paper proposes a preference-conditioned Bellman operator based on Chebyshev scalarization, embedded within an end-to-end framework for learning deterministic Pareto optimal policies. The proposed operator possesses an envelope property that guarantees the value function’s upper bound encompasses the true Pareto front and ensures monotonic convergence. Consequently, it enables the synthesis of approximately Pareto optimal policies under arbitrary user-specified preferences. Experimental results demonstrate that the method effectively recovers complex trade-offs among multiple objectives and achieves comprehensive coverage of the Pareto front.
Existing nonlinear scalarization methods struggle to guarantee uniqueness and continuity of the mapping from preferences to Pareto-optimal solutions, thereby limiting coverage of dense Pareto fronts. This work proposes Smooth Chebyshev Scalarization and rigorously proves that the induced Pareto-optimal return vectors are uniquely determined by preferences and Lipschitz continuous with respect to them. Building upon an occupancy measure formulation, the authors develop the Concave Mirror Descent Policy Iteration (CMDPI) algorithm and establish its equivalence to a KL-regularized MDP, ensuring policy continuity in preference space. Integrated with a KL-regularized deep actor-critic architecture, the proposed method achieves the best average hypervolume ranking across eight MO-Gymnasium tasks and significantly outperforms discrete-action variants in continuous control settings.
This work addresses the challenge in dynamic environments where existing reinforcement learning methods rely on manually specified reward weights, hindering their ability to adaptively balance primary objective optimization with constraint satisfaction. The authors propose MAMO, a novel approach that, for the first time, formulates reward weight selection as a learnable task. By leveraging multi-agent reinforcement learning, MAMO decouples task execution from objective design. The method integrates Lagrangian-inspired reward shaping with online policy learning to automatically adjust the trade-off between cost minimization and performance constraints. Evaluated in non-stationary settings, MAMO significantly enhances policy adaptability, robustness, and adherence to constraints without requiring manual tuning of reward coefficients.
Real-world decision-making often involves multiple conflicting objectives, and the true reward function is typically difficult to specify a priori. To address this challenge, this work proposes the LEMUR framework, which extends preference-based reinforcement learning to the multi-objective setting for the first time. LEMUR leverages preference feedback from multiple human evaluators to jointly learn, in an end-to-end manner, both a multi-objective policy and its corresponding reward model, without requiring any pre-defined reward functions. By integrating multi-objective reward modeling with policy optimization, the method achieves a significant performance advantage over existing baselines across several benchmark tasks, effectively balancing and optimizing among competing objectives in an unsupervised setting.
This work addresses the challenge of recovering a complete Pareto front policy from multiple Pareto-optimal expert demonstrations while avoiding dominated solutions caused by aggregating conflicting state-action trajectories. The authors propose Multi-output Augmented Behavioral Cloning (MA-BC), an algorithm that separates divergent state-action pairs and integrates only non-conflicting data to enable efficient multi-objective imitation learning. They establish, for the first time, a minimax lower bound for this problem and prove that MA-BC achieves a faster statistical convergence rate, matching the theoretical optimum. Empirical evaluations on both discrete environments and continuous linear quadratic regulator (LQR) tasks demonstrate that MA-BC significantly outperforms approaches that independently learn individual expert policies, with experimental results aligning closely with the theoretical guarantees.
This work addresses the issue of premature convergence in multi-agent multi-objective optimization, which often arises from behavioral homogenization. To mitigate this, the study introduces a behavioral entropy maximization mechanism into multi-objective evolutionary algorithms for the first time. Specifically, within the NSGA-II framework, it integrates policy entropy rewards with multi-objective fitness evaluation to explicitly promote behavioral diversity while preserving Pareto optimality. This approach effectively alleviates behavioral collapse and substantially enhances exploration capability. Experimental results in the rover domain demonstrate that, compared to the NSGA-II baseline, the proposed method achieves up to a 48% improvement in hypervolume metric, along with significantly enhanced solution set quality and diversity.