🤖 AI Summary
This study addresses the failure to recover optimal rankings in pairwise comparison sorting under heterogeneous preferences due to insufficient repeated sampling. Based on the Bradley-Terry model, we investigate the minimum number of repeated comparisons required to identify the item with the highest expected utility. Through maximum likelihood estimation and theoretical analysis, we establish the optimality of a logarithmic number of repeated comparisons. Furthermore, we propose a Russian roulette-based randomized algorithm that reduces the expected number of comparisons to a constant level. This approach compresses the sample complexity from polynomial to logarithmic order. Experiments on both synthetic and Arena semi-synthetic datasets validate the effectiveness of our method across varying degrees of preference heterogeneity.
📝 Abstract
We study ranking models by population-average utility from pairwise comparisons when preferences vary across users and tasks. Prior work shows that a single comparison per user can be insufficient to identify the alternative with the highest average utility, even with arbitrarily many users (Golz et al., 2025). We investigate how many repeated comparisons within each user-task context are necessary and sufficient for ranking recovery. Under a heterogeneous Bradley-Terry model with fixed inverse temperature, we start with a naive MLE-based algorithm that requires $Ω(1/Δ^2)$ repeated comparisons per context to ensure ranking recovery. We then present two MLE-based variants and a randomized Russian Roulette-style algorithm that recover the ranking using $O(\log(1/Δ))$ repeated comparisons per context, and we prove that this logarithmic dependence is optimal. Despite this worst-case requirement, our Russian Roulette algorithm uses only $O(1)$ comparisons per context in expectation. Synthetic experiments and semi-synthetic experiments based on Arena data compare the four algorithms in settings with varying levels of preference heterogeneity and under varying context distributions.