🤖 AI Summary
This study investigates whether version upgrades of large language models enhance the accuracy of probabilistic estimates in prediction markets. To this end, it introduces the RMF benchmark, comprising 3,000 resolved binary questions, and evaluates the Claude and Qwen model families under a zero-shot protocol through hierarchical statistical analyses incorporating Brier scores, Murphy decomposition, empirical baseline comparisons, and paired hypothesis testing. The findings reveal that model-tier upgrades exert non-uniform effects on predictive accuracy. Specifically, while Claude outperforms empirical baselines, iterative version updates yield no significant gains, whereas smaller Qwen variants exhibit inadequate calibration. This work demonstrates that interpreting predictive scores necessitates accounting for baselines and question composition, thereby offering a systematic framework for evaluating the probabilistic calibration capabilities of large language models.
📝 Abstract
We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protocol. We assess Brier scores relative to an empirical base-rate predictor, examine their Murphy decomposition, and compare models through paired differences, with results stratified by event category and timing relative to training cutoffs. In the reported post-cutoff stratum, the four Claude models achieve Brier scores of 0.183-0.192, improving on the base-rate reference by 0.024-0.033. Qwen 32B does not significantly outperform that reference, although its paired Brier is 0.024 lower than that of the 7B checkpoint. The evaluated Claude version and tier upgrades yield no significant improvement. Within-model differences across event categories exceed the observed differences among Claude variants. These results show why forecasting scores should be interpreted alongside simple probability baselines and question composition: under this protocol, newer versions or higher model tiers do not consistently produce more accurate probabilities.