🤖 AI Summary
This work addresses the instability or failure of maximum likelihood estimation in the Bradley–Terry model when the comparison graph is disconnected or nearly separable. To mitigate this issue, the authors propose two interpretable data augmentation–based regularization strategies: introducing pseudo-matches between all pairs of players and incorporating virtual players with fixed strengths. The former yields shrinkage estimates of ability parameters, while the latter simultaneously resolves the inherent non-identifiability due to location invariance in a natural manner. Experiments on 2025 Major League Baseball season data demonstrate that, with appropriate tuning, these regularized approaches accurately replicate the effects of ridge regression while preserving intuitive semantic interpretations grounded in the augmented data.
📝 Abstract
Paired comparison models are useful for estimating latent abilities or preferences from binary outcomes, but maximum likelihood estimation can be unstable or fail when the comparison graph is disconnected or nearly separated. Ridge regularization addresses these difficulties by shrinking ability parameters toward a common center, but it can obscure the simple likelihood interpretation that makes Bradley-Terry and Thurstone-Mosteller models attractive to practitioners. This paper describes two data-augmentation perspectives on regularization. The first adds fractional pseudo-games between every pair of competitors. The second adds a fixed-strength phantom player and gives each real competitor a weighted pseudo-win and pseudo-loss against that player. Both approaches yield finite, shrunken estimates; the phantom-player construction also resolves the usual location nonidentifiability without an explicit linear constraint. For the Bradley-Terry model, the two augmentations lead to transparent penalty functions that can be compared directly with ridge penalties. An application to the 2025 Major League Baseball regular season illustrates that tuned pseudo-game and phantom-player regularization can closely reproduce ridge-regularized strength estimates while retaining an intuitive augmented-data representation.