diverse subset selection

Designs and analyzes algorithms and procedures for choosing a small, informative or representative subset from a larger pool that optimize objectives such as diversity, informativeness, risk exposure, or task relevance. This work covers greedy and top‑k selection heuristics, risk‑aware and selective‑prediction criteria, constrained optimization formulations, and scalable implementations for large candidate sets.

diversesubsetselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In evolutionary multi-objective optimization, indicator-based subset selection suffers from high computational overhead in local search. To address this, we propose a dual-candidate-list acceleration strategy: for the first time, neighborhood constraints are introduced into this task, constructing both a nearest-neighbor and a random candidate list, coupled with a serialized switching mechanism to balance search efficiency and solution quality—effectively mitigating challenges posed by Pareto front discontinuities. Our method integrates k-nearest-neighbor candidate set construction, random sampling, and indicator-driven local search (e.g., hypervolume optimization). Experiments demonstrate speedups of several-fold to over an order of magnitude on continuous fronts, with negligible degradation in subset quality; on discontinuous fronts, it significantly enhances robustness and solution quality.

Addresses discontinuity in Pareto fronts with candidate listsImproves efficiency of subset selection in multi-objective optimizationReduces computational cost in local search

Pareto Optimization with Robust Evaluation for Noisy Subset Selection

Jan 12, 2025
YX
Yi-Heng Xu
🏛️ Nanjing University

This paper addresses the noisy subset selection problem under cardinality constraints, where objective function evaluations are highly noisy and computationally expensive. We propose the first method that integrates a robust evaluation function into a multi-objective Pareto optimization framework, simultaneously optimizing solution quality and subset size. Our approach jointly incorporates noise modeling, noise-resilient sampling, and evolutionary strategies to balance robustness and efficiency. Experiments on real-world influence maximization and sparse regression benchmarks demonstrate significant improvements over greedy algorithms, POSS, and PONSS. Ablation studies confirm that the robust evaluation module is the key driver of performance gains.

Computational EfficiencyNoise MitigationSubset Selection

Analyzing the Landscape of the Indicator-based Subset Selection Problem

Apr 11, 2025
KK
Keisuke Korogi
🏛️ Yokohama National University

This work investigates the fitness landscape characteristics of the Indicator-based Subset Selection Problem (ISSP), aiming to uncover how quality indicator types and Pareto front geometry jointly shape landscape structure. Methodologically, it introduces the first systematic integration of classical landscape analysis—such as fitness–distance correlation and ruggedness—with precise Local Optima Networks (LONs) to construct an interpretable ISSP landscape characterization framework. Key findings reveal that the ε-indicator induces pronounced neutrality plateaus alongside high-density multimodal local optima, highlighting the landscape’s extreme sensitivity to both indicator choice and front geometry. These results provide theoretical foundations for algorithm design in subset selection and advance the interpretability of ISSP performance analysis.

Analyzing ISSP impact from quality indicators and Pareto frontExploring high neutrality and local optima in ISSP instancesUnderstanding ISSP landscape for better subset selection methods

This paper studies the fair top-$k$ selection problem: ensuring representative inclusion of minority or historically disadvantaged groups when selecting the top $k$ items from high-dimensional data using a linear scoring function. We first establish the inherent computational hardness of this problem in high dimensions by proving its NP-hardness. For small $k$, we propose an efficient algorithm with theoretical guarantees and its parallel implementation; for large $k$, we design a scalable, practical surrogate method. Our approach integrates linear weighted modeling, rigorous complexity analysis, and hardware-aware optimization targeting multi-core CPUs and GPUs. Extensive experiments on real-world datasets demonstrate that our methods achieve speedups of several orders of magnitude over state-of-the-art baselines, significantly improving both efficiency and scalability for fair top-$k$ selection in large-scale, high-dimensional settings.

Developing scalable algorithms for large, high-dimensional datasets.Ensuring selected subsets represent minority groups fairly.Finding fair linear scoring functions for top-k selection.

Selecting the Best Optimizing System

Jan 09, 2022
NS
Nian Si
🏛️ Hong Kong University of Science and Technology | University of California, Berkeley

This paper addresses the Stochastic Black-Box Optimization Selection (SBOS) problem: identifying, among multiple stochastic systems with continuous decision variables, the system whose optimal decision yields the best expected performance—without prior knowledge and under a finite sampling budget. We propose the first formal SBOS framework that jointly integrates intra-system optimization via stochastic gradient descent and inter-system comparison via sequential elimination, enabling synergistic optimization across both levels. We theoretically establish that the probability of incorrect selection converges exponentially with the sampling budget. Empirical evaluation across three real-world SBOS scenarios demonstrates that our method significantly reduces the misselection probability and maintains robust superiority across varying budget sizes and problem dimensions.

Designing algorithms to sequentially eliminate inferior systems under budget constraintsIntegrating gradient descent with elimination methods to minimize false selection probabilitySelecting the best system among finite contenders with optimal decisions

Latest Papers

What's happening recently
View more

Picking a Representative Set of Solutions in Multiobjective Optimization: Axioms, Algorithms, and Experiments

Nov 13, 2025
NB
Niclas Boehmer
🏛️ Hasso Plattner Institute | University of Potsdam

In multi-objective optimization, the Pareto-optimal solution set is often excessively large, imposing heavy cognitive burden on decision-makers. Method: This paper proposes “Directional Coverage,” a novel representativeness metric, and conducts axiomatic analysis within a multi-winner voting framework to expose counterintuitive behaviors of existing indicators and characterize their computational complexity boundaries under varying objective structures. The approach integrates axiomatic modeling, computational complexity theory, and empirical evaluation to systematically compare how diverse quality metrics affect solution set representativeness. Results: Experiments demonstrate that Directional Coverage achieves superior or comparable performance to state-of-the-art metrics in diversity, convergence, and directional sensitivity. Crucially, the choice of quality metric fundamentally determines representativeness outcomes—providing both theoretical foundations and practical tools for Pareto set reduction.

Analyzing quality measures for multiobjective optimization through axiomatic studyDeveloping new directed coverage measure and computational complexity analysisSelecting representative Pareto optimal solutions to reduce decision maker cognitive load

This work addresses the computational challenges of solving large-scale convex mixed-integer quadratic programs (MIQPs), which arise in applications such as subset portfolio selection and become particularly difficult when the covariance matrix has a high condition number or weight constraints are tight. To tackle this, the authors propose DASH, a novel method that introduces a decreasing active-set hierarchy for dimensionality reduction in MIQP for the first time. DASH leverages active-set analysis to reduce problem dimensionality and integrates seamlessly with commercial solvers like Gurobi to enhance optimization efficiency. Experimental results demonstrate that DASH significantly outperforms Gurobi alone on a range of challenging portfolio instances, with solution quality improvements positively correlated with problem difficulty, thereby accelerating convergence and yielding higher-quality optimal solutions.

Dimensionality ReductionMixed Integer Quadratic ProgrammingNP-hard

This study addresses the challenge of efficiently generating and managing Pareto-optimal solution sets (SOS) in heterogeneous multi-task environments. It proposes an evolutionary multi-task optimization framework to construct compact, task-specific SOS repositories for real-world applications such as engineering design, inventory management, and hyperparameter optimization. The work introduces a novel similarity metric between Pareto sets and, for the first time, systematically validates the cross-domain applicability of SOS. Through visualization and objective space analysis, it reveals dynamic patterns in solution set performance across diverse task contexts. Experimental results demonstrate that the proposed approach effectively captures inter-task differences in solution sets and significantly enhances decision-making support across varying scenarios.

evolutionary multitaskingmultiobjective optimizationmultitask optimization

This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.

Bayesian inferencebest subset selectionexpert knowledge

MODE: Multi-Objective Adaptive Coreset Selection

Dec 24, 2025
TM
Tanmoy Mukherjee
🏛️ Université d’Artois

Conventional static core-set selection fails to adapt to the heterogeneous requirements across different training stages. Method: This paper proposes a dynamic multi-objective adaptive core-set selection framework that dynamically switches sampling strategies according to training progression—emphasizing class balance in early stages, feature diversity in mid-stages, and prediction uncertainty in late stages—thereby enabling the first training-process-aware, multi-objective co-optimization. Contribution/Results: We theoretically establish a (1−1/e)-approximation guarantee. By integrating submodular optimization, active learning, and representation analysis, our method achieves O(n log n) computational efficiency. Empirically, it attains full-dataset accuracy on multiple benchmarks while significantly reducing memory overhead. Moreover, it is the first work to quantitatively characterize the dynamic evolution of data utility throughout training.

Adapts selection criteria to different training phasesDynamically combines coreset selection strategies for model performanceReduces memory requirements while maintaining competitive accuracy

Hot Scholars

PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
ST

Shu-Tao Xia

SIGS, Tsinghua University
coding and information theorymachine learningcomputer visionAI security
YL

Yadan Luo

ARC DECRA and Senior Lecturer, University of Queensland
Generalization3D VisionAutonomous Driving
XZ

Xiawu Zheng

Associate Professor, IEEE Senior Member, Xiamen University
Automated Machine LearningNetwork CompressionNeural Architecture SearchAutoML
FR

Feng Ruan

Department of EECS, University of California, Berkeley
Machine LearningStatistics