Score
Designs, builds, and analyzes composite objective structures—such as fitness, utility, or reward functions and their scalarizations—that combine and balance two or more criteria to guide optimization and evaluation of models, policies, or algorithms. Implements and tunes weighted or joint objectives, multi-objective RL reward schemes, and multi-metric evaluation protocols to quantify trade-offs, prevent single-metric failure modes, and steer convergence toward desired balanced behaviors.
Multi-objective reinforcement learning (MORL) suffers from non-unique mappings between policy parameter space and multi-objective performance space, poor interpretability, and low efficiency in Pareto frontier search. To address these challenges, this paper proposes an interpretable MORL framework based on local linear mapping—first embedding bidirectional parameter–performance interpretability into algorithm design. Specifically, it models the parameter-to-performance mapping locally as linear, enabling real-time semantic interpretation of policy objectives; supports zero-shot cross-domain policy transfer without retraining; and facilitates interpretable gradient-guided approximation of the Pareto frontier. Evaluated on multiple benchmark tasks, our method achieves significant improvements: +18.7% in Pareto frontier coverage and 2.3× acceleration in convergence speed, while outperforming state-of-the-art methods in both explanation quality and search efficiency.
Highly constrained multi-objective optimization problems (MOOPs) with repetitive fitness landscapes—such as the multi-objective vehicle routing problem with time windows—pose significant computational challenges due to the high cost of repeatedly applying expensive multi-objective algorithms across similar instances. Method: This paper proposes a transfer-based goal programming framework: first, a representative instance is solved using a costly multi-objective algorithm to obtain a high-quality approximate Pareto set; this set is then embedded into a goal programming model to construct three objective-specific single-objective fitness functions, which guide efficient single-objective algorithms to rapidly solve subsequent similar instances. Contribution/Results: This work is the first to systematically exploit inter-instance fitness landscape similarity for knowledge transfer in MOOPs, balancing solution quality and computational efficiency. Experiments demonstrate that the approach significantly reduces runtime while generating high-quality compromise solutions, effectively reconciling effectiveness and efficiency.
Existing multi-objective optimization research predominantly focuses on conflicting objectives and Pareto fronts, overlooking the prevalent “aligned objectives” scenario in machine learning—where objectives are non-conflicting and mutually reinforcing. This work formally defines the aligned multi-objective optimization problem and breaks from the traditional Pareto paradigm by proposing the first gradient-based optimization framework tailored to this setting. Methodologically, it introduces a dynamic weight allocation and gradient normalization fusion algorithm grounded in gradient direction alignment analysis, accompanied by theoretical convergence guarantees. Compared to naive strategies such as weighted sum, the approach achieves significantly improved optimization efficiency and stability. Empirical evaluation on multi-task learning and large language model training demonstrates synchronous performance gains across all objectives, faster convergence, enhanced robustness, and scalability to large-scale, highly correlated objective sets.
This work addresses the challenge of efficiently approximating the Pareto front in multi-objective optimization (MOO). We propose a novel set-based optimization framework grounded in the R2 utility function, which reformulates MOO as a single-objective optimization problem over solution sets. We theoretically establish that the R2 utility is both monotonic and submodular—properties enabling a greedy algorithm with a guaranteed (1−1/e) approximation ratio. Integrating this with Bayesian optimization, our approach achieves efficient, high-coverage approximation of the Pareto front. Methodologically, it unifies scalarization, submodular optimization, and Bayesian optimization within a coherent framework. Empirical evaluation on multi-objective Bayesian optimization benchmarks demonstrates significant improvements in convergence speed and Pareto front quality, validating both its practical efficacy and theoretical advantages.
The R2 indicator for bi-objective optimization lacks strict Pareto compatibility—i.e., adding a dominated solution may not increase the indicator value. Method: This paper proposes an analytical variant of R2 based on a continuous uniform Tchebycheff utility function. We theoretically prove that this continuous R2 exhibits strict Pareto compatibility for bi-objective problems: adding any non-dominated solution strictly increases the indicator, and adding any dominated solution necessarily increases it. We further devise an exact O(N log N) algorithm, enabling the first efficient and theoretically compliant unary quality assessment. Results: Experiments show that the proposed indicator achieves evaluation performance comparable to hypervolume (HV), yet with significantly higher computational efficiency. It thus fills a critical gap in bi-objective set-quality indicators by simultaneously offering rigorous theoretical guarantees and practical scalability.
This work proposes a novel extension of reinforcement learning from AI feedback (RLAIF) to multi-objective adaptive traffic signal control, addressing the common limitation of existing approaches that prioritize a single dominant objective due to conflicting goals and fail to accommodate diverse user preferences. By leveraging large language models to generate multi-objective preference labels—without requiring manual annotation or intricate reward function engineering—the method trains adaptive control policies that reflect varying user priorities. Experimental results demonstrate that the proposed approach effectively balances competing objectives such as traffic efficiency, fairness, and energy consumption, significantly enhancing both the practicality and scalability of traffic signal control strategies in real-world urban environments.
This work addresses a critical limitation in multi-objective reinforcement learning, where policy evaluation based solely on value vectors often overlooks behavioral differences, leading to ambiguous decision-making. To resolve this, the authors propose an exploratory diagnostic framework that explicitly incorporates behavioral divergence into Pareto front analysis for the first time. By integrating trajectory clustering with visualization techniques, the method quantitatively reveals the behavioral diversity among Pareto-optimal policies. Empirical validation on both grid-world and continuous control benchmarks demonstrates its effectiveness: even in complex tasks, the framework clearly delineates behavioral distinctions between policies, thereby offering decision-makers a richer, more informative basis for policy selection.
This work addresses the challenge that non-experts face in systematically constructing alignment reward functions that reflect human preferences. The authors propose a three-step framework: first translating natural language objectives into measurable outcome variables, then modeling the selection of reward terms as a minimum-cost partial set cover problem grounded in a causal graph, and finally iteratively fitting linear reward weights through preference queries. This approach is the first to deterministically characterize the conflict-free feasible weight region and provably converges to a target accuracy within \(O(n \log \kappa)\) queries. By integrating causal reasoning, max-flow computation, and convex feasibility solving, the framework achieves high efficiency, interpretability, and theoretical guarantees, substantially lowering the barrier for non-experts to design aligned reward functions.
This work addresses the challenge in multi-objective reinforcement learning (MORL) of generalizing to arbitrary preference weightings when user preferences are unknown. It introduces reward-free reinforcement learning (RFRL) into MORL for the first time in a systematic manner, leveraging RFRL objectives as auxiliary tasks to enhance policy generalization to unseen preference-weighted rewards. The proposed approach integrates a conditional policy network with a preference-guided exploration mechanism, enabling effective knowledge transfer across reward functions and sample-efficient learning. Evaluated on multiple MO-Gymnasium benchmarks, the method significantly outperforms existing MORL algorithms, achieving state-of-the-art performance in both final policy quality and data efficiency.
This work addresses the challenge in multi-objective reinforcement learning (MORL) where single-policy approaches often fail to fully recover the Pareto front due to gradient interference and policy representation collapse. To overcome this, the authors propose the D³PO framework, which decouples the optimization of individual objectives to preserve distinct learning signals, delays preference fusion, and introduces a scaled diversity regularizer to enhance policy sensitivity to preferences. As the first method to systematically identify and mitigate gradient interference and representation collapse in preference-conditioned policies, D³PO achieves state-of-the-art or comparable performance using only a single deployable policy across multiple MORL benchmarks, significantly improving both the coverage and quality of the recovered Pareto front as measured by hypervolume and expected utility metrics.