sequential subset selection

Designs and analyzes algorithms and procedures that choose subsets of items or observations sequentially (iteratively or online) to optimize an explicit objective such as information, utility, or predictive performance under constraints. Work includes developing greedy and optimal subset‑selection methods, mechanisms to control subset size versus computation, and strategies to account for dependencies among elements when constructing informative subsets.

sequentialsubsetselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Solving the Best Subset Selection Problem via Suboptimal Algorithms

Mar 31, 2025
VS
Vikram Singh
🏛️ University of Central Oklahoma | The University of Alabama

The NP-hard problem of best-subset selection in high-dimensional linear regression motivates this work. We propose an efficient, suboptimal algorithmic framework that integrates greedy search, regularized path tracking, and cross-validation-based model evaluation to yield a stable and scalable solution pipeline. Compared with mainstream heuristic approaches—including LASSO and orthogonal matching pursuit (OMP)—our method substantially reduces computational cost in ultra-high-dimensional settings (p ≥ 1000) while achieving superior trade-offs between model sparsity and predictive accuracy. Comprehensive benchmark experiments on both synthetic data and diverse real-world datasets demonstrate that our approach consistently attains higher solution quality and greater robustness than state-of-the-art baselines. By bridging efficiency and statistical reliability, the proposed framework establishes a new paradigm for high-dimensional sparse modeling.

Addresses computational challenges in best subset selectionCompares performance of new and existing methodsProposes suboptimal algorithms for high-dimensional data

Greedy Selection under Independent Increments: A Toy Model Analysis

Jun 22, 2025
HY
Huitao Yang
🏛️ Fudan University

This paper studies a class of multi-stage iterative selection problems: given $N$ independent and identically distributed discrete-time stochastic processes with independent increments, one retains a fixed number of processes at each stage, aiming to maximize the probability of selecting the process achieving the global maximum value. We rigorously prove that, under the independent-increments assumption, a greedy policy—retaining at each stage the processes with the largest current observed values—achieves global optimality; i.e., it is equivalent to the optimal stopping policy. This result challenges the conventional intuition that greedy strategies are suboptimal in multi-stage selection, providing the first theoretically optimal heuristic for elimination-based sequential decision-making (e.g., online hiring, resource scheduling) under non-Markovian dynamics. Methodologically, we integrate probabilistic analysis with optimal stopping theory. Although the result relies on strong independence assumptions, it establishes an extensible theoretical foundation for high-dimensional or approximately independent settings.

Justify greedy heuristics in multi-stage eliminationProve greedy selection maximizes final valueStudy iterative selection of i.i.d. processes

Pareto Optimization with Robust Evaluation for Noisy Subset Selection

Jan 12, 2025
YX
Yi-Heng Xu
🏛️ Nanjing University

This paper addresses the noisy subset selection problem under cardinality constraints, where objective function evaluations are highly noisy and computationally expensive. We propose the first method that integrates a robust evaluation function into a multi-objective Pareto optimization framework, simultaneously optimizing solution quality and subset size. Our approach jointly incorporates noise modeling, noise-resilient sampling, and evolutionary strategies to balance robustness and efficiency. Experiments on real-world influence maximization and sparse regression benchmarks demonstrate significant improvements over greedy algorithms, POSS, and PONSS. Ablation studies confirm that the robust evaluation module is the key driver of performance gains.

Computational EfficiencyNoise MitigationSubset Selection

Combinatorial Selection with Costly Information

Dec 05, 2024
SC
Shuchi Chawla
🏛️ University of Texas at Austin | Cornell University

This paper studies sequential stochastic combinatorial optimization with costly information acquisition: it jointly optimizes solution quality and observation cost within an acyclic Markov decision process (MDP) framework, under matroid constraints and bandit-style feedback. Addressing the limitation of prior work—restricted to special cases—we propose a novel cost-allocation bound coupled with a local approximation framework, enabling lossless approximate composition of solutions for arbitrary component MDPs. Our approach transcends structural restrictions inherent in classical models such as Pandora’s Box, and unifies treatment of a broad class of variants—including the newly introduced Weighing Scale problem—yielding constant-factor optimal approximation algorithms either for maximizing expected reward or minimizing total cost.

Approximate solutions for bandit superprocesses with matroid constraintsNew approximations for combinatorial Pandora's Box variantsOptimizing stochastic variables with costly information acquisition

Selecting the Best Optimizing System

Jan 09, 2022
NS
Nian Si
🏛️ Hong Kong University of Science and Technology | University of California, Berkeley

This paper addresses the Stochastic Black-Box Optimization Selection (SBOS) problem: identifying, among multiple stochastic systems with continuous decision variables, the system whose optimal decision yields the best expected performance—without prior knowledge and under a finite sampling budget. We propose the first formal SBOS framework that jointly integrates intra-system optimization via stochastic gradient descent and inter-system comparison via sequential elimination, enabling synergistic optimization across both levels. We theoretically establish that the probability of incorrect selection converges exponentially with the sampling budget. Empirical evaluation across three real-world SBOS scenarios demonstrates that our method significantly reduces the misselection probability and maintains robust superiority across varying budget sizes and problem dimensions.

Designing algorithms to sequentially eliminate inferior systems under budget constraintsIntegrating gradient descent with elimination methods to minimize false selection probabilitySelecting the best system among finite contenders with optimal decisions

Latest Papers

What's happening recently
View more

Removal of Redundant Candidate Points for the Exact D-Optimal Design Problem

Aug 31, 2025
RH
Radoslav Harman
🏛️ Comenius University

Computing exact D-optimal designs over large finite candidate sets is computationally prohibitive due to excessive memory and time requirements. Method: We propose a redundancy elimination method grounded in necessary conditions for approximate designs. It integrates convex optimization to compute an approximate design, integer-constrained screening to identify candidate support points, and mixed-integer second-order cone programming to recover the exact optimal solution. Contribution/Results: We establish, for the first time, a theoretical link between approximate and exact designs—proving that, asymptotically, the optimal support set is contained within the maximum-variance subset. This enables stepwise, order-of-magnitude compression of the candidate set. Our approach reduces candidate sets with tens of millions of points by several orders of magnitude, solves previously intractable large-scale instances, and delivers solutions with certified global optimality.

Eliminating redundant points without losing optimality guaranteesEnabling exact design computation via mixed-integer optimizationReducing large candidate sets for exact D-optimal designs

This work addresses the fundamental challenge of efficiently selecting an informative subset of $n$ samples from a large dataset of size $N$ for parameter estimation, particularly when data volume is massive or labeling costs are prohibitive. Building upon optimal approximate design theory, the authors propose the first general-purpose subdata selection framework that is theoretically convergent and accommodates multiple optimality criteria. They develop efficient algorithms capable of approximating the optimal solution across arbitrary $N$ and $n$. The selected subdata achieve information efficiency nearly matching the theoretical upper bound, substantially outperforming existing methods. Moreover, this study provides the first tight upper and lower bounds to rigorously evaluate the efficiency of any subset selection strategy.

information-based samplingNP-hardoptimal design

This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.

familywise errorhypothesis configurationmultiple hypotheses

Near Optimal Inference for the Best-Performing Algorithm

Aug 07, 2025
AP
Amichai Painsky
🏛️ Tel Aviv University

This work addresses the problem of reliably identifying the smallest subset of machine learning algorithms most likely to contain the optimal algorithm on unseen datasets, given performance observations on a limited set of benchmark datasets—particularly when performance differences are marginal and high-confidence guarantees are required. We propose a novel subset selection framework grounded in multinomial statistical inference, which ensures that the true optimal algorithm lies within the selected subset with provable confidence. The method is computationally efficient both asymptotically and in finite-sample regimes, and we provide the first proof establishing that its sample complexity matches the information-theoretic lower bound. Theoretical analysis and empirical evaluation demonstrate that our approach significantly improves the joint trade-off between subset minimality and confidence level compared to existing methods, enabling more accurate and compact identification of the optimal algorithm candidates.

Identify best-performing algorithm from benchmark datasetsImprove subset selection methods with theoretical boundsSelect minimal subset including top algorithm confidently

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
JS

Jingbo Shang

Associate Professor, UC San Diego
Natural Language ProcessingData MiningDeep LearningInformation Extraction
SW

Shaobo Wang

Shanghai Jiao Tong University
Large Language ModelData-Centric AIData SynthesisData Selection
ZW

Zichen Wen

Shanghai Jiao Tong University
Efficient AITrustworthy AILarge Language ModelMachine Learning
KL

Kaixin Li

National University of Singapore
Machine LearningNatural Language ProcessingCode IntelligenceGUI Agents