Iterative Hypothesis Pruning and Distribution-based Early Labeling for Sequential Hypothesis Testing

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the problem of efficiently identifying the true hypothesis under limited samples for an active decision-maker. We propose a deterministic multi-stage hypothesis elimination algorithm that, at each stage, selects actions maximizing distributional distinguishability—measured by total variation distance (TVD)—to guide sequential sampling. Crucially, we couple hypothesis clustering with iterative pruning, enabling batch elimination of hypotheses within similarity-defined clusters. To our knowledge, this is the first work to establish a joint clustering–sampling optimization framework grounded in distributional similarity, where adaptive sampling is guided by a TVD-based minimax criterion. We derive necessary and sufficient conditions for asymptotic error convergence to zero. Theoretically, the algorithm achieves asymptotically optimal sample complexity, admits a provable error bound, and runs in polynomial time—substantially improving convergence efficiency in large hypothesis spaces.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchMachine Learning: Active LearningReasoning under Uncertainty: Stochastic Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web text
📝 Abstract
We consider the problem where an active Decision-Maker (DM) is tasked to identify the true hypothesis using as few samples as possible while maintaining accuracy. The DM collects samples according to its determined actions and knows the distributions under each hypothesis. We propose the $Φ$-$Δ$ algorithm, a deterministic and adaptive multi-stage hypothesis-elimination algorithm where the DM selects an action, applies it repeatedly, and discards hypotheses in light of its obtained samples. The DM selects actions based on maximal separation expressed by the maximal minimal Total Variation Distance (TVD) between each two possible output distributions. To further optimize the search (in terms of the mean number of samples required to separate hypotheses), close distributions (in TVD) are clustered, and the algorithm eliminates whole clusters rather than individual hypotheses. We extensively analyze our algorithm and show it is asymptotically optimal as the desired error probability approaches zero. Our analysis also includes identifying instances when the algorithm is asymptotically optimal in the number of hypotheses, bounding the mean number of samples per-stage and in total, characterizing necessary and sufficient conditions for vanishing error rates when clustering hypotheses, evaluating algorithm complexity, and discussing its optimality in finite regimes.
Problem

Research questions and friction points this paper is trying to address.

Minimizing sample usage while maintaining hypothesis testing accuracy
Developing adaptive algorithm for sequential elimination of hypotheses
Optimizing action selection through distribution clustering and separation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Iterative hypothesis pruning reduces candidate hypotheses
Clustering distributions by TVD distance for efficiency
Adaptive multi-stage elimination with maximal separation actions
💼 Related Jobs
No related jobs found.
G
George Vershinin
The School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel
A
Asaf Cohen
The School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel
O
Omer Gurewitz
The School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel