Score
Designs and fits probabilistic pairwise-comparison models that infer latent item quality scores from noisy pairwise votes while explicitly estimating judge-specific bias parameters and covariate effects; builds Bayesian Bradley–Terry–style likelihoods with hierarchical shrinkage priors to regularize bias estimates and aggregate biased comparisons.
In generative model evaluation, human pairwise comparisons exhibit higher consistency than single-item ratings; however, existing Bradley–Terry–based methods often neglect annotator quality heterogeneity and lack theoretical convergence guarantees, undermining robustness and interpretability. To address this, we propose BBQ (Bayesian Bradley–Terry with Quality-aware Raters), the first Bayesian Bradley–Terry framework that explicitly models annotator quality. We establish a theoretical guarantee of monotonic likelihood convergence. BBQ employs an EM algorithm for rater-aware Bayesian inference, adaptively weighting noisy annotations based on inferred annotator reliability. Experiments demonstrate that BBQ achieves faster convergence and superior uncertainty calibration. Crucially, it maintains stable and reliable ranking and scoring under high noise and crowdsourced settings. By jointly inferring item utilities and annotator quality, BBQ significantly enhances the robustness and interpretability of human preference modeling in generative evaluation.
This work addresses the limitations of the classical Bradley–Terry model, which assumes transitive preferences and thus fails to capture cyclic intransitivities prevalent in real-world competitive networks, leading to biased estimates and miscalibrated uncertainty quantification. To overcome this, we propose a Bayesian intransitive Bradley–Terry model that, for the first time, integrates combinatorial Hodge decomposition into the pairwise comparison framework. This approach explicitly decomposes pairwise preferences into a transitive gradient flow and an intransitive curl flow, with adaptive regularization achieved through a global–local shrinkage prior. The model naturally reduces to the classical Bradley–Terry formulation in the absence of intransitivity and enables uncertainty quantification at the triplet level. Experiments demonstrate that our method significantly outperforms existing Bayesian intransitive approaches in estimation accuracy, uncertainty calibration, and computational efficiency, while effectively identifying cyclic competitive advantages.
Statistical inference in large-scale pairwise comparison settings becomes challenging as the number of subjects grows unbounded. Method: This paper establishes a unified asymptotic theory framework by characterizing the Fisher information matrix as a weighted graph Laplacian and conducting refined spectral analysis. Contribution/Results: It derives near-optimal asymptotic normality and individual convergence rate $O(1/n)$ for the maximum likelihood estimator—applicable beyond the Bradley–Terry model to a broad class of generalized pairwise comparison models. Crucially, the framework enables cross-model unified analysis, eliminating the need for model-specific derivations. Extensive validation on synthetic data and real-world professional tennis match outcomes confirms high estimation accuracy and validity of hypothesis tests. The theory provides a scalable, interpretable statistical foundation for high-dimensional ordinal data analysis.
The Bradley–Terry (BT) model underpins pairwise comparison ranking, yet its theoretical foundations remain fragmented across disparate motivations, lacking a unified statistical interpretation. Method: This paper systematically unifies over ten independent derivations—including maximum likelihood estimation, random utility theory, Elo-style dynamical evolution, game-theoretic equilibrium analysis, and Bayesian inference—and introduces two novel perspectives: a gamified motivation framework and a progressive probabilistic interpretation. Contribution/Results: The analysis reveals the BT model as a “minimally structured preference encoder,” elucidating its fundamental role in learning-to-rank through rigorous statistical modeling and asymptotic analysis. This unified characterization establishes a principled foundation for enhancing algorithmic interpretability, designing robust ranking systems, and enabling cross-domain transfer—thereby bridging theoretical understanding with practical deployment in preference learning.
To address the challenge of learning total orders from large-scale, sparse, and highly noisy pairwise comparison data, this paper proposes a generative, noise-aware ranking algorithm. Methodologically, it explicitly incorporates pairwise comparison confidence into a scalable ranking framework for the first time; introduces a quasiconvex approximation strategy for the underlying nonconvex optimization problem, solved via iterative reweighted minimization combined with the Primal-Dual Hybrid Gradient method; and extends the Bradley–Terry model to capture heterogeneous noise. Empirically, the method achieves a 0.1 improvement in Kendall tau over state-of-the-art approaches, maintains robustness under 10% erroneous comparisons, and reduces computational time by an order of magnitude—enabling ranking in seconds. Validation on real-world datasets demonstrates significant superiority over active learning baselines. Overall, the approach establishes a new Pareto-optimal trade-off between accuracy and efficiency in large-scale noisy ranking.
This study addresses the susceptibility of parametric models to misspecification in high-dimensional paired comparison data by proposing a semiparametric modeling framework that incorporates an unspecified latent distribution to capture item merits and covariate effects. The method employs kernel-based least squares estimation and, for the first time under a diverging dimensionality setting, enables semiparametric analysis of paired comparisons with covariates, balancing model flexibility with statistical inferential validity. Theoretical results establish the consistency and asymptotic normality of the proposed estimator. Extensive simulations and an empirical analysis of NBA data demonstrate the effectiveness and practical utility of the approach.
This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.
本文提出共识校准方法,通过分治策略独立校准各时段并重构后验分布,解决了大规模、稀疏且频繁更新的IRT题库的高效校准问题。
This work addresses the non-reciprocity in pairwise comparison data arising jointly from genuine scale variations and random noise. The authors propose an additive decomposition model that disentangles the observed comparison matrix into three components: a consistent non-reciprocal structure encoding a global ranking, a symmetric component capturing scale-induced discrepancies, and Gaussian noise representing judgment errors. Within a probabilistic framework, the method distinguishes structural non-reciprocity from stochastic perturbations, thereby preserving symmetric information essential for analyzing scale effects while avoiding the information loss incurred by enforcing reciprocity. The resulting admissible ranking region enables noise calibration and validation of scale plausibility, substantially outperforming conventional approaches that project observations directly onto reciprocal matrices.
This study investigates whether pairwise comparisons genuinely reflect model accuracy or are confounded by stylistic preferences and annotator biases. The authors reformulate five standard benchmarks as open-ended generation tasks and integrate pairwise comparisons, Elo rating aggregation, and Spearman correlation analysis with causal inference techniques. Their systematic evaluation—conducted in settings where ground-truth labels are available—demonstrates for the first time that Elo rankings exhibit strong alignment with accuracy-based rankings (Spearman correlation > 0.9), significantly outperforming direct scoring, especially under weak annotator conditions. Furthermore, they identify “answer repetition at the end” as a key causal factor driving human preference, while finding that stylistic variation and annotator bias exert only minimal influence on overall model rankings.