Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of reproducible benchmarks and training environments for LLM-based forecasting agents by proposing the first replayable forecasting benchmark framework. By integrating Polymarket data with 18.8 million time-stamped news articles, the framework constructs a simulation environment that supports continuous historical backtracking, research, and forecast revision, enabling iterative evaluation and feedback learning without awaiting new events. Experimental results demonstrate that research tools significantly reduce Brier scores across various models and that supervised fine-tuning (SFT) serves as an effective proof of concept; however, model performance remains inferior to historical market consensus. Overall, this work establishes a standardized paradigm for the systematic evaluation and data-driven optimization of forecasting agents.
📝 Abstract
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.
Problem

Research questions and friction points this paper is trying to address.

LLM forecasting agents
benchmarking
prediction markets
replayable environments
agent training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Replayable Environments
LLM Forecasting Agents
Prediction Markets
Benchmarking
Supervised Fine-Tuning
🔎 Similar Papers