Score
Designs, builds, and analyzes ensembles of models by selecting and generating candidate ensemble methods and members (e.g., parallel or heterogeneous ensembles), choosing aggregation, voting or decision rules, and determining weighting and calibration strategies. Implements and evaluates ensemble aggregation and simulation procedures to compare predictive and probabilistic performance, validate ensemble choices on held-out data, and derive best-practice combination and evaluation methods.
Time-series ensemble forecasting faces a critical trade-off between predictive accuracy and computational cost. This paper systematically evaluates ten base models and eight ensemble strategies on the M5 and VN1 retail datasets, measuring performance in point forecasting (RMSE) and probabilistic forecasting (CRPS), alongside computational overhead. Methodologically, we analyze ensemble size scalability, propose an “efficiency-driven ensemble” paradigm, and assess downsampling-based retraining frequency reduction. Key contributions: (1) Ensembles of only two to three models achieve near-optimal accuracy; (2) The efficiency-driven paradigm reduces average computational cost by over 40% while retaining ≥95% of baseline accuracy; (3) Reducing retraining frequency cuts training overhead by up to 70%, with negligible impact on point forecasts and robust performance in probabilistic forecasting. Results confirm that ensembling consistently improves prediction—especially probabilistic calibration—but high accuracy typically incurs high cost. Our framework delivers a scalable, cost-effective ensemble strategy for resource-constrained deployment.
This work addresses three key challenges in extreme weather forecasting: insufficient sampling of tail distributions, biased uncertainty quantification, and inadequate representation of internal variability. To this end, we propose HENS—a hyper-scale ensemble forecasting framework comprising 7,424 members. Methodologically, HENS is the first to integrate spherical Fourier neural operators (SFNO), bred-vector initial-condition perturbations, and multi-checkpoint model-parameter perturbations on a ten-million-node spherical grid, enabling a high-fidelity, low-overhead parallel ensemble generation system. Compared with operational systems, HENS achieves comparable physical consistency while reducing computational cost by several orders of magnitude. It significantly improves tail-sampling accuracy for 4σ extreme events, enhances single-member skill and trajectory coverage, and reduces ensemble outlier rates. Moreover, HENS provides a more comprehensive characterization of internal variability, concurrently improving both grid-scale forecast coverage and reliability.
Predicting low-probability, high-impact extreme weather events under climate change remains challenging due to the computational expense and limited ensemble size of traditional numerical weather prediction (NWP) models. Method: This work introduces the first application of the Spherical Fourier Neural Operator (SFNO) to ultra-large-ensemble forecasting—supporting 1,000–10,000 members—by integrating perturbed-parameter ensembles (to represent model uncertainty) and bred-vector initial-condition perturbations (to represent initial-state uncertainty). A multi-scale spectral diagnostic framework and extreme-event indices are incorporated to ensure physical consistency and forecast reliability. Contribution/Results: Experiments confirm spectral stability over time—validated via key spectral tests—and calibrated probabilistic forecasts of extremes exhibit discrimination and reliability comparable to ECMWF’s IFS. This establishes a scalable, physics-informed AI paradigm for modeling long-tail climate risks.
Diffusion models for weather forecasting suffer from autoregressive rollout, high computational cost, and error accumulation at high temporal resolution. To address these issues, this paper proposes Continuous Ensemble Diffusion (CED), a parallel ensemble forecasting framework based on continuous-time modeling. CED treats meteorological fields as a continuous-time stochastic process governed by an implicit ordinary differential equation (ODE), enabling parallel generation of temporally coherent ensemble trajectories in a single forward pass—eliminating autoregressive iteration entirely. The method integrates implicit ODE solvers, ensemble probabilistic calibration, and parallel denoising, supporting arbitrary temporal resolution and flexible hybridization with conventional rollout. On global weather forecasting benchmarks, CED achieves state-of-the-art performance: it improves skill scores significantly, reduces the Continuous Ranked Probability Score (CRPS) by 12.3%, and exhibits superior spread-error consistency.
This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
This study investigates optimal parallel heterogeneous ensemble strategies for small- to medium-scale tabular classification tasks. Through systematic experiments on 56 OpenML CC18 datasets, the authors find that the instability of Blending and Stacking is mutually independent, and that Robust Soft Voting significantly outperforms Hard Voting in multi-class settings. Building on these insights, they propose a robust ensemble strategy and validate its effectiveness on the TabArena platform. The recommended approach significantly surpasses the Single Best method across 28 additional tasks and matches or exceeds the performance of all individual ensemble techniques considered.
This work proposes a novel ensemble framework termed Behavioral Profiling Ensemble, which addresses the limitation of traditional ensemble methods that assign static weights based on model diversity and validation sets, thereby ignoring the varying capabilities of base models across different regions of the input space. The key innovation lies in introducing the concept of “intrinsic behavioral profiles” for each base model, capturing its characteristic response patterns. During inference, ensemble weights are dynamically allocated according to the deviation between the test instance and each model’s behavioral profile. Notably, the approach eliminates the need for validation sets or assumptions about model diversity. Extensive experiments on multiple synthetic and real-world datasets demonstrate that the proposed method consistently outperforms existing ensemble techniques, achieving substantial improvements in predictive accuracy, computational efficiency, and memory overhead.
Quantifying the contribution of individual submodels to the overall predictive performance of ensemble systems is crucial for enhancing interpretability and construction efficiency. This work proposes and implements a unified R package that, for the first time in the R ecosystem, systematically supports model importance assessment across diverse ensemble methods under both point and probabilistic forecasting frameworks, with full compatibility with the hubverse infrastructure. The package offers flexible importance metrics and robust handling of missing values, substantially improving the understanding of submodel roles. It thereby empowers researchers to efficiently construct, diagnose, and optimize ensemble forecasting systems.
This work addresses the overfitting or underfitting issues in surrogate models caused by improper smoothness design by proposing an adaptive ensemble surrogate-assisted evolutionary algorithm. For the first time, smoothness is explicitly formulated as a multi-objective optimization problem that jointly minimizes approximation error and model complexity. This approach automatically optimizes the structure of radial basis function networks to construct a robust ensemble surrogate model with diverse smoothness levels and incorporates a multi-model collaborative pre-screening mechanism. Extensive experiments on both single-objective benchmark functions and real-world computationally expensive optimization problems demonstrate that the proposed method significantly outperforms state-of-the-art surrogate-assisted evolutionary algorithms, with statistically significant advantages.