🤖 AI Summary
Existing rating systems treat draws as equivalent to half-wins and half-losses, ignoring the empirically observed nonlinear increase in draw probability with player strength—particularly pronounced in strategic games like chess—leading to biased strength estimation. This paper proposes the first Bayesian dynamic rating framework that explicitly embeds a strength-dependent draw mechanism, abandoning the conventional linear simplification. Our method enables efficient online inference via closed-form posterior updates and a single-step Newton–Raphson approximation. Experiments on large-scale correspondence chess data from the International Correspondence Chess Federation demonstrate significantly improved model fit, more stable strength tracking over time, and superior long-term predictive accuracy. The core contribution is the first formal modeling of the draw-generation process as a nonlinear function of player strength, seamlessly integrated into a principled Bayesian dynamic rating paradigm.
📝 Abstract
Competitor rating systems for head-to-head games are typically used to measure playing strength from game outcomes. Ratings computed from these systems are often used to select top competitors for elite events, for pairing players of similar strength in online gaming, and for players to track their own strength over time. Most implemented rating systems assume only win/loss outcomes, and treat occurrences of ties as the equivalent to half a win and half a loss. However, in games such as chess, the probability of a tie (draw) is demonstrably higher for stronger players than for weaker players, so that rating systems ignoring this aspect of game results may produce strength estimates that are unreliable. We develop a new rating system for head-to-head games based on a model by Glickman (2025) that explicitly acknowledges that a tie may depend on the strengths of the competitors. The approach uses a Bayesian dynamic modeling framework. Within each time period, posterior updates are computed in closed form using a single Newton-Raphson iteration evaluated at the prior mean. The approach is demonstrated on a large dataset of chess games played in International Correspondence Chess Federation tournaments.