Score
Designs and analyzes algorithms and procedures for choosing a small, informative or representative subset from a larger pool that optimize objectives such as diversity, informativeness, risk exposure, or task relevance. This work covers greedy and top‑k selection heuristics, risk‑aware and selective‑prediction criteria, constrained optimization formulations, and scalable implementations for large candidate sets.
In evolutionary multi-objective optimization, indicator-based subset selection suffers from high computational overhead in local search. To address this, we propose a dual-candidate-list acceleration strategy: for the first time, neighborhood constraints are introduced into this task, constructing both a nearest-neighbor and a random candidate list, coupled with a serialized switching mechanism to balance search efficiency and solution quality—effectively mitigating challenges posed by Pareto front discontinuities. Our method integrates k-nearest-neighbor candidate set construction, random sampling, and indicator-driven local search (e.g., hypervolume optimization). Experiments demonstrate speedups of several-fold to over an order of magnitude on continuous fronts, with negligible degradation in subset quality; on discontinuous fronts, it significantly enhances robustness and solution quality.
This paper addresses the noisy subset selection problem under cardinality constraints, where objective function evaluations are highly noisy and computationally expensive. We propose the first method that integrates a robust evaluation function into a multi-objective Pareto optimization framework, simultaneously optimizing solution quality and subset size. Our approach jointly incorporates noise modeling, noise-resilient sampling, and evolutionary strategies to balance robustness and efficiency. Experiments on real-world influence maximization and sparse regression benchmarks demonstrate significant improvements over greedy algorithms, POSS, and PONSS. Ablation studies confirm that the robust evaluation module is the key driver of performance gains.
This work investigates the fitness landscape characteristics of the Indicator-based Subset Selection Problem (ISSP), aiming to uncover how quality indicator types and Pareto front geometry jointly shape landscape structure. Methodologically, it introduces the first systematic integration of classical landscape analysis—such as fitness–distance correlation and ruggedness—with precise Local Optima Networks (LONs) to construct an interpretable ISSP landscape characterization framework. Key findings reveal that the ε-indicator induces pronounced neutrality plateaus alongside high-density multimodal local optima, highlighting the landscape’s extreme sensitivity to both indicator choice and front geometry. These results provide theoretical foundations for algorithm design in subset selection and advance the interpretability of ISSP performance analysis.
This paper studies the fair top-$k$ selection problem: ensuring representative inclusion of minority or historically disadvantaged groups when selecting the top $k$ items from high-dimensional data using a linear scoring function. We first establish the inherent computational hardness of this problem in high dimensions by proving its NP-hardness. For small $k$, we propose an efficient algorithm with theoretical guarantees and its parallel implementation; for large $k$, we design a scalable, practical surrogate method. Our approach integrates linear weighted modeling, rigorous complexity analysis, and hardware-aware optimization targeting multi-core CPUs and GPUs. Extensive experiments on real-world datasets demonstrate that our methods achieve speedups of several orders of magnitude over state-of-the-art baselines, significantly improving both efficiency and scalability for fair top-$k$ selection in large-scale, high-dimensional settings.
This paper addresses the Stochastic Black-Box Optimization Selection (SBOS) problem: identifying, among multiple stochastic systems with continuous decision variables, the system whose optimal decision yields the best expected performance—without prior knowledge and under a finite sampling budget. We propose the first formal SBOS framework that jointly integrates intra-system optimization via stochastic gradient descent and inter-system comparison via sequential elimination, enabling synergistic optimization across both levels. We theoretically establish that the probability of incorrect selection converges exponentially with the sampling budget. Empirical evaluation across three real-world SBOS scenarios demonstrates that our method significantly reduces the misselection probability and maintains robust superiority across varying budget sizes and problem dimensions.
In multi-objective optimization, the Pareto-optimal solution set is often excessively large, imposing heavy cognitive burden on decision-makers. Method: This paper proposes “Directional Coverage,” a novel representativeness metric, and conducts axiomatic analysis within a multi-winner voting framework to expose counterintuitive behaviors of existing indicators and characterize their computational complexity boundaries under varying objective structures. The approach integrates axiomatic modeling, computational complexity theory, and empirical evaluation to systematically compare how diverse quality metrics affect solution set representativeness. Results: Experiments demonstrate that Directional Coverage achieves superior or comparable performance to state-of-the-art metrics in diversity, convergence, and directional sensitivity. Crucially, the choice of quality metric fundamentally determines representativeness outcomes—providing both theoretical foundations and practical tools for Pareto set reduction.
This work addresses the computational challenges of solving large-scale convex mixed-integer quadratic programs (MIQPs), which arise in applications such as subset portfolio selection and become particularly difficult when the covariance matrix has a high condition number or weight constraints are tight. To tackle this, the authors propose DASH, a novel method that introduces a decreasing active-set hierarchy for dimensionality reduction in MIQP for the first time. DASH leverages active-set analysis to reduce problem dimensionality and integrates seamlessly with commercial solvers like Gurobi to enhance optimization efficiency. Experimental results demonstrate that DASH significantly outperforms Gurobi alone on a range of challenging portfolio instances, with solution quality improvements positively correlated with problem difficulty, thereby accelerating convergence and yielding higher-quality optimal solutions.
This study addresses the challenge of efficiently generating and managing Pareto-optimal solution sets (SOS) in heterogeneous multi-task environments. It proposes an evolutionary multi-task optimization framework to construct compact, task-specific SOS repositories for real-world applications such as engineering design, inventory management, and hyperparameter optimization. The work introduces a novel similarity metric between Pareto sets and, for the first time, systematically validates the cross-domain applicability of SOS. Through visualization and objective space analysis, it reveals dynamic patterns in solution set performance across diverse task contexts. Experimental results demonstrate that the proposed approach effectively captures inter-task differences in solution sets and significantly enhances decision-making support across varying scenarios.
This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.
Conventional static core-set selection fails to adapt to the heterogeneous requirements across different training stages. Method: This paper proposes a dynamic multi-objective adaptive core-set selection framework that dynamically switches sampling strategies according to training progression—emphasizing class balance in early stages, feature diversity in mid-stages, and prediction uncertainty in late stages—thereby enabling the first training-process-aware, multi-objective co-optimization. Contribution/Results: We theoretically establish a (1−1/e)-approximation guarantee. By integrating submodular optimization, active learning, and representation analysis, our method achieves O(n log n) computational efficiency. Empirically, it attains full-dataset accuracy on multiple benchmarks while significantly reducing memory overhead. Moreover, it is the first work to quantitatively characterize the dynamic evolution of data utility throughout training.