Score
Designs and implements methods to set, learn, and adapt scalar weights that combine multiple objectives in a decision-making or learning system, including algorithms that tune objective coefficients from feedback and integrate weight updates into planners or policies. Analyzes how weight choices affect trade-offs and system behavior to align outputs with stated or inferred preferences and to adjust weights as intents or environments change.
This study systematically exposes fundamental limitations of scalarization-based methods in multi-objective reinforcement learning (MORL) with discrete action and observation spaces: poor Pareto-front coverage, low robustness, and strong dependence on environmental properties and front geometry. To address these issues, we propose an inner-loop multi-policy architecture and comparatively evaluate three representative approaches—linear scalarization, Chebyshev scalarization, and non-scalarized Pareto Q-learning—under both outer-loop single-policy and inner-loop multi-policy paradigms. Results demonstrate that Pareto Q-learning significantly improves solution-set diversity and stability over scalarized methods. Moreover, the inner-loop multi-policy design effectively mitigates scalarization’s sensitivity to weight selection and susceptibility to local optima. Our empirical analysis provides both a novel methodological paradigm and rigorous evidence for developing robust, scalable MORL algorithms.
This paper addresses the challenge of feature weight assignment in discrete multi-objective data analysis. We propose a dynamic feature weighting method grounded in evolutionary game theory, modeling feature weights as population states on the standard simplex and employing analytically tractable replicator dynamics to iteratively evolve weights over a normalized data matrix. We rigorously prove global convergence to a unique non-degenerate interior equilibrium, thereby eliminating the weight collapse commonly observed in conventional approaches. Our key contribution is the first application of an evolutionary game-theoretic framework to feature weighting—yielding a method that guarantees theoretical convergence, offers intuitive interpretability, and ensures numerical stability. This establishes a novel paradigm for multi-objective feature selection, bridging rigorous mathematical foundations with practical applicability.
This work addresses the challenge of efficiently approximating the Pareto front in multi-objective optimization (MOO). We propose a novel set-based optimization framework grounded in the R2 utility function, which reformulates MOO as a single-objective optimization problem over solution sets. We theoretically establish that the R2 utility is both monotonic and submodular—properties enabling a greedy algorithm with a guaranteed (1−1/e) approximation ratio. Integrating this with Bayesian optimization, our approach achieves efficient, high-coverage approximation of the Pareto front. Methodologically, it unifies scalarization, submodular optimization, and Bayesian optimization within a coherent framework. Empirical evaluation on multi-objective Bayesian optimization benchmarks demonstrates significant improvements in convergence speed and Pareto front quality, validating both its practical efficacy and theoretical advantages.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This paper addresses optimal policy assignment under partial treatment coverage in heterogeneous populations, where experiments only implement a subset of possible treatment values, limiting generalizability to untested interventions. Method: We propose the first framework integrating shape-constrained partial identification of treatment effects—incorporating monotonicity or convexity constraints—with a minimax regret criterion, formulated as a tractable mixed-integer linear program (MILP). The method combines nonparametric conditional average treatment effect (CATE) estimation, shape restrictions, minimax optimization, and efficient linear/integer programming solvers. Contribution/Results: Applied to a Kenyan rural electricity subsidy experiment, our framework recommends novel, experimentally untested treatment levels—covering nearly the entire population—while reducing maximum regret by over 60%. It substantially enhances policy extrapolation capability and robustness beyond conventional methods constrained to observed treatment supports.
This work addresses a critical limitation in multi-objective reinforcement learning, where policy evaluation based solely on value vectors often overlooks behavioral differences, leading to ambiguous decision-making. To resolve this, the authors propose an exploratory diagnostic framework that explicitly incorporates behavioral divergence into Pareto front analysis for the first time. By integrating trajectory clustering with visualization techniques, the method quantitatively reveals the behavioral diversity among Pareto-optimal policies. Empirical validation on both grid-world and continuous control benchmarks demonstrates its effectiveness: even in complex tasks, the framework clearly delineates behavioral distinctions between policies, thereby offering decision-makers a richer, more informative basis for policy selection.
This study addresses the challenge of efficiently generating and managing Pareto-optimal solution sets (SOS) in heterogeneous multi-task environments. It proposes an evolutionary multi-task optimization framework to construct compact, task-specific SOS repositories for real-world applications such as engineering design, inventory management, and hyperparameter optimization. The work introduces a novel similarity metric between Pareto sets and, for the first time, systematically validates the cross-domain applicability of SOS. Through visualization and objective space analysis, it reveals dynamic patterns in solution set performance across diverse task contexts. Experimental results demonstrate that the proposed approach effectively captures inter-task differences in solution sets and significantly enhances decision-making support across varying scenarios.
This study addresses the challenge of optimizing time-varying weight sequences in Double Linear Policy (DLP) by formulating it as a receding-horizon stochastic optimal control problem. The approach dynamically maximizes risk-adjusted returns subject to survivability and expected positive return constraints. A key contribution is the first derivation of an analytical gradient for this non-convex objective, which is embedded within a stochastic model predictive control (SMPC) framework and solved via the L-BFGS-B algorithm to enable closed-loop, dynamic optimization of DLP weights. Empirical results demonstrate that the proposed method significantly outperforms both fixed-weight and predetermined time-varying strategies in terms of risk-adjusted returns and drawdown control.
This work addresses the challenge in multi-objective optimization where conventional uniform sampling of scalarization weights often yields uneven coverage of the Pareto front (PF), failing to align with diverse user preferences. The study establishes, for the first time, an analytical relationship between the geometric traversal speed of scalarization paths along the PF and the underlying weight distribution. Building on this insight, it proposes a PF-aware sampling strategy based on the arc-length cumulative distribution function (CDF) and its inverse mapping, which can be implemented either analytically or iteratively. The method guarantees uniform PF coverage and exhibits linear convergence. Empirical evaluations across bi-objective bandit problems, multi-objective Gymnasium environments, and large language model alignment tasks demonstrate significant performance gains over existing approaches, offering both strong theoretical guarantees and practical efficacy.