🤖 AI Summary
This study addresses the problem of efficient sequential data collection in political polling, aiming to terminate sampling as early as possible while preserving statistical validity. To this end, the authors propose a posterior optimal sequential testing procedure based on e-values, establishing—for the first time—a statistical testing framework for the Schulze voting system and devising tailored tests for Condorcet and Borda rules. The key innovation lies in solving the challenging construction of optimal e-values under non-convex composite hypotheses with multivariate Bernoulli data, achieved via an efficient Frank–Wolfe algorithm for computing reverse information projection, applicable to general composite null and alternative hypotheses. Empirical evaluations on both synthetic simulations and real-world data from the 2022 French presidential election demonstrate that the proposed method substantially outperforms state-of-the-art approaches in both statistical power and sample efficiency.
📝 Abstract
We are interested in conducting political polls sequentially, so that one can stop acquiring data as soon as possible while safely yielding statistically significant results. Building off e-values, which have recently become a useful tool to create sequential testing methods, we develop a theory of posterior optimal e-values. We use voting as a convenient example on which to illustrate our method.
First, we design statistical tests for Condorcet and Borda voting system, and also for Schulze voting system which we are the first to tackle statistically. Then, we study the construction of optimal sequential e-values in the deceptively simple setting of multivariate Bernoulli data, with general composite null and alternative hypothesis sets $\mathcal{H}_0$ and $\mathcal{H}_1$. We give a way to compute these e-values using an efficient Frank-Wolfe algorithm, giving a pretty general way to compute Reverse Information Projections, even when $\mathcal{H}_0$ corresponds to a non-convex parameter set. Finally, we illustrate the efficiency, both in terms of power and sample size of our method. We compare with state of the art in both simulated and real data experiments, with application to French 2022 presidential election data.