LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of large language models (LLMs) predominantly rely on static, retrospective benchmarks, which fail to capture their predictive capabilities under real-world uncertainty. To address this gap, this work introduces the first prospective, real-time evaluation protocol and open-source platform tailored for unresolved events. The framework employs a factorial experimental design to systematically assess LLM performance across varying information access modalities, prompting strategies, and prediction horizons, while integrating timestamp logging, pattern validation, tool-use tracing, and cost tracking. Experiments on 104 matches and 15 tournament-related questions from the 2026 FIFA World Cup reveal that models with web access only marginally outperform those without (a Brier score improvement of 0.023), highlighting the current limitations of LLMs in forecasting real-world outcomes.
📝 Abstract
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.
Problem

Research questions and friction points this paper is trying to address.

large language models
forecasting
real-world predictions
sports
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

prospective benchmarking
real-world forecasting
factorial design
LLM evaluation
sports prediction
🔎 Similar Papers
No similar papers found.