Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge that high costs and restricted access to top-tier large language models (LLMs) hinder high-quality event forecasting. To overcome this limitation, this work proposes a weak-to-strong aggregation paradigm that leverages linear pooling to learn optimal weights for combining multiple weaker LLMs into collaborative predictions. The method’s performance is evaluated using the Brier score on the ForecastBench benchmark. Experimental results demonstrate that the proposed approach significantly enhances both predictive accuracy and calibration without relying on the strongest individual model. Notably, in 11 out of 16 experimental configurations, the aggregated weak models matched or surpassed the performance of the best single model, with all configuration errors remaining within a 5% margin.
πŸ“ Abstract
Large language models (LLMs) are increasingly used to forecast real-world events, but access to the strongest individual forecaster may be costly or otherwise constrained. We study weak-to-strong forecast aggregation: can individually weaker LLM forecasters be aggregated to outperform a stronger forecaster? Using ForecastBench (Karger et al., 2025), we evaluate 70 LLM forecasters across 16 comparison groups, each with more than 1,000 shared subquestions, yielding 1,121 weaker-model pairs. Within each group, we identify the strongest individual by test Brier score and evaluate aggregates composed exclusively of weaker forecasters, with aggregation weights learned on separate training data. We find substantial evidence of weak-to-strong improvement. Learned linear pooling identifies a weaker pair that matches or outperforms the strongest individual in 11 of 16 groups and comes within 5% of its Brier score in all 16 groups. We also find that these improvements do not rely on having a near-best constituent and are generally accompanied by good calibration. Additional analyses show that adding more models does not consistently improve performance, and competitive weaker-model aggregates also remain available under practical constraints.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Forecast Aggregation
Weak-to-Strong
Prediction
Brier Score
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weak-to-Strong Aggregation
Learned Linear Pooling
LLM Forecasting
Calibration
ForecastBench
πŸ”Ž Similar Papers
No similar papers found.