Beyond forecast leaderboards: Measuring individual model importance based on contribution to ensemble accuracy

📅 2024-12-12
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.

Technology Category

Machine Learning: Ensemble MethodsNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsReasoning under Uncertainty: Relational Probabilistic Models

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Ensemble forecasts often outperform forecasts from individual standalone models, and have been used to support decision-making and policy planning in various fields. As collaborative forecasting efforts to create effective ensembles grow, so does interest in understanding individual models' relative importance in the ensemble. To this end, we propose two practical methods that measure the difference between ensemble performance when a given model is or is not included in the ensemble: a leave-one-model-out algorithm and a leave-all-subsets-of-models-out algorithm, which is based on the Shapley value. We explore the relationship between these metrics, forecast accuracy, and the similarity of errors, both analytically and through simulations. We illustrate this measure of the value a component model adds to an ensemble in the presence of other models using US COVID-19 death forecasts. This study offers valuable insight into individual models' unique features within an ensemble, which standard accuracy metrics alone cannot reveal.
Problem

Research questions and friction points this paper is trying to address.

Measuring individual model importance in ensemble forecasting accuracy
Proposing practical methods to assess model contribution to ensemble performance
Revealing unique model features beyond standard accuracy metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leave-one-model-out algorithm for ensemble contribution
Shapley value-based leave-all-subsets-out algorithm
Measuring model importance through ensemble accuracy impact
🔎 Similar Papers
University of Massachusetts
M
Minsu Kim
School of Public Health and Health Sciences, University of Massachusetts, 715 North Pleasant Street, Amherst, 01003, Massachusetts, United States of America
E
E. Ray
School of Public Health and Health Sciences, University of Massachusetts, 715 North Pleasant Street, Amherst, 01003, Massachusetts, United States of America
N
N. Reich
School of Public Health and Health Sciences, University of Massachusetts, 715 North Pleasant Street, Amherst, 01003, Massachusetts, United States of America