bias-aware pairwise ranking

Designs and fits probabilistic pairwise-comparison models that infer latent item quality scores from noisy pairwise votes while explicitly estimating judge-specific bias parameters and covariate effects; builds Bayesian Bradley–Terry–style likelihoods with hierarchical shrinkage priors to regularize bias estimates and aggregate biased comparisons.

bias-awarepairwiseranking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Efficient Bayesian Inference from Noisy Pairwise Comparisons

Oct 10, 2025
TA
Till Aczel
🏛️ ETH Zürich | Mabyduck

In generative model evaluation, human pairwise comparisons exhibit higher consistency than single-item ratings; however, existing Bradley–Terry–based methods often neglect annotator quality heterogeneity and lack theoretical convergence guarantees, undermining robustness and interpretability. To address this, we propose BBQ (Bayesian Bradley–Terry with Quality-aware Raters), the first Bayesian Bradley–Terry framework that explicitly models annotator quality. We establish a theoretical guarantee of monotonic likelihood convergence. BBQ employs an EM algorithm for rater-aware Bayesian inference, adaptively weighting noisy annotations based on inferred annotator reliability. Experiments demonstrate that BBQ achieves faster convergence and superior uncertainty calibration. Crucially, it maintains stable and reliable ranking and scoring under high noise and crowdsourced settings. By jointly inferring item utilities and annotator quality, BBQ significantly enhances the robustness and interpretability of human preference modeling in generative evaluation.

Aggregating inconsistent ratings from unreliable human evaluatorsEvaluating generative models with noisy human pairwise comparisonsImproving Bradley-Terry models with rater quality and convergence guarantees

This work addresses the limitations of the classical Bradley–Terry model, which assumes transitive preferences and thus fails to capture cyclic intransitivities prevalent in real-world competitive networks, leading to biased estimates and miscalibrated uncertainty quantification. To overcome this, we propose a Bayesian intransitive Bradley–Terry model that, for the first time, integrates combinatorial Hodge decomposition into the pairwise comparison framework. This approach explicitly decomposes pairwise preferences into a transitive gradient flow and an intransitive curl flow, with adaptive regularization achieved through a global–local shrinkage prior. The model naturally reduces to the classical Bradley–Terry formulation in the absence of intransitivity and enables uncertainty quantification at the triplet level. Experiments demonstrate that our method significantly outperforms existing Bayesian intransitive approaches in estimation accuracy, uncertainty calibration, and computational efficiency, while effectively identifying cyclic competitive advantages.

Bradley-Terry modelcompetitive networkscycle effects

Statistical inference for pairwise comparison models

Jan 16, 2024
RH
Ruijian Han
🏛️ The Hong Kong Polytechnic University | University of Alberta | University of Kentucky

Statistical inference in large-scale pairwise comparison settings becomes challenging as the number of subjects grows unbounded. Method: This paper establishes a unified asymptotic theory framework by characterizing the Fisher information matrix as a weighted graph Laplacian and conducting refined spectral analysis. Contribution/Results: It derives near-optimal asymptotic normality and individual convergence rate $O(1/n)$ for the maximum likelihood estimator—applicable beyond the Bradley–Terry model to a broad class of generalized pairwise comparison models. Crucially, the framework enables cross-model unified analysis, eliminating the need for model-specific derivations. Extensive validation on synthetic data and real-world professional tennis match outcomes confirms high estimation accuracy and validity of hypothesis tests. The theory provides a scalable, interpretable statistical foundation for high-dimensional ordinal data analysis.

Addresses statistical inference with diverging subject counts in pairwise comparison modelsEstablishes asymptotic normality for maximum likelihood estimators in pairwise comparison modelsProvides theoretical foundations for inference beyond the Bradley-Terry model

The many routes to the ubiquitous Bradley-Terry model

Dec 21, 2023
IH
Ian Hamilton
🏛️ University of Warwick

The Bradley–Terry (BT) model underpins pairwise comparison ranking, yet its theoretical foundations remain fragmented across disparate motivations, lacking a unified statistical interpretation. Method: This paper systematically unifies over ten independent derivations—including maximum likelihood estimation, random utility theory, Elo-style dynamical evolution, game-theoretic equilibrium analysis, and Bayesian inference—and introduces two novel perspectives: a gamified motivation framework and a progressive probabilistic interpretation. Contribution/Results: The analysis reveals the BT model as a “minimally structured preference encoder,” elucidating its fundamental role in learning-to-rank through rigorous statistical modeling and asymptotic analysis. This unified characterization establishes a principled foundation for enhancing algorithmic interpretability, designing robust ranking systems, and enabling cross-domain transfer—thereby bridging theoretical understanding with practical deployment in preference learning.

Compare various approaches to item ratingExplore motivations for using Bradley-Terry modelPresent novel and known statistical model derivations

Ranking with Confidence for Large Scale Comparison Data

Feb 03, 2022
FV
Filipa Valdeira
🏛️ NOVA School of Science and Technology (NOVA FCT)

To address the challenge of learning total orders from large-scale, sparse, and highly noisy pairwise comparison data, this paper proposes a generative, noise-aware ranking algorithm. Methodologically, it explicitly incorporates pairwise comparison confidence into a scalable ranking framework for the first time; introduces a quasiconvex approximation strategy for the underlying nonconvex optimization problem, solved via iterative reweighted minimization combined with the Primal-Dual Hybrid Gradient method; and extends the Bradley–Terry model to capture heterogeneous noise. Empirically, the method achieves a 0.1 improvement in Kendall tau over state-of-the-art approaches, maintains robustness under 10% erroneous comparisons, and reduces computational time by an order of magnitude—enabling ranking in seconds. Validation on real-world datasets demonstrates significant superiority over active learning baselines. Overall, the approach establishes a new Pareto-optimal trade-off between accuracy and efficiency in large-scale noisy ranking.

Confidence EstimationLarge-scale DataRanking Algorithm

Latest Papers

What's happening recently
View more

This study addresses the susceptibility of parametric models to misspecification in high-dimensional paired comparison data by proposing a semiparametric modeling framework that incorporates an unspecified latent distribution to capture item merits and covariate effects. The method employs kernel-based least squares estimation and, for the first time under a diverging dimensionality setting, enables semiparametric analysis of paired comparisons with covariates, balancing model flexibility with statistical inferential validity. Theoretical results establish the consistency and asymptotic normality of the proposed estimator. Extensive simulations and an empirical analysis of NBA data demonstrate the effectiveness and practical utility of the approach.

covariate effectshigh-dimensional inferencemodel misspecification

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

This work addresses the non-reciprocity in pairwise comparison data arising jointly from genuine scale variations and random noise. The authors propose an additive decomposition model that disentangles the observed comparison matrix into three components: a consistent non-reciprocal structure encoding a global ranking, a symmetric component capturing scale-induced discrepancies, and Gaussian noise representing judgment errors. Within a probabilistic framework, the method distinguishes structural non-reciprocity from stochastic perturbations, thereby preserving symmetric information essential for analyzing scale effects while avoiding the information loss incurred by enforcing reciprocity. The resulting admissible ranking region enables noise calibration and validation of scale plausibility, substantially outperforming conventional approaches that project observations directly onto reciprocal matrices.

Admissible ranking regionsDecision analysisNoise calibration

This study investigates whether pairwise comparisons genuinely reflect model accuracy or are confounded by stylistic preferences and annotator biases. The authors reformulate five standard benchmarks as open-ended generation tasks and integrate pairwise comparisons, Elo rating aggregation, and Spearman correlation analysis with causal inference techniques. Their systematic evaluation—conducted in settings where ground-truth labels are available—demonstrates for the first time that Elo rankings exhibit strong alignment with accuracy-based rankings (Spearman correlation > 0.9), significantly outperforming direct scoring, especially under weak annotator conditions. Furthermore, they identify “answer repetition at the end” as a key causal factor driving human preference, while finding that stylistic variation and annotator bias exert only minimal influence on overall model rankings.

accuracy rankingsElo rankinggenerative models

Hot Scholars

FM

Fred Morstatter

University of Southern California, Information Sciences Institute
Social Media MiningData ScienceData MiningMachine Learning
CS

Chuan Shi

Beijing University of Posts and Telecommunications
data miningmachine learningsocial network analysis
RD

Rebecca Dorn

University of Southern California
AI FairnessNatural Language ProcessingComputational Social Science
KL

Kristina Lerman

Professor of Informatics, Indiana University
Social NetworksData ScienceArtificial IntelligenceComputational Social Science
CB

Charles Bickham

PhD Candidate at the University of Southern California
Human-Centered AI Mental Health Computational Social Science Social Media Data Science