Score
Designs and implements methods that select which models or forecasters to include in an ensemble and how to allocate a fixed sampling or evaluation budget across them, then constructs the combined ensemble (via selection and/or weighting) to maximize predictive accuracy or another performance metric subject to resource constraints. Analyzes and optimizes trade-offs between budget allocation, model selection, and ensemble performance, producing algorithms and procedures for sampling, allocation, and ensemble optimization under cost or sample-size limits.
Time-series ensemble forecasting faces a critical trade-off between predictive accuracy and computational cost. This paper systematically evaluates ten base models and eight ensemble strategies on the M5 and VN1 retail datasets, measuring performance in point forecasting (RMSE) and probabilistic forecasting (CRPS), alongside computational overhead. Methodologically, we analyze ensemble size scalability, propose an “efficiency-driven ensemble” paradigm, and assess downsampling-based retraining frequency reduction. Key contributions: (1) Ensembles of only two to three models achieve near-optimal accuracy; (2) The efficiency-driven paradigm reduces average computational cost by over 40% while retaining ≥95% of baseline accuracy; (3) Reducing retraining frequency cuts training overhead by up to 70%, with negligible impact on point forecasts and robust performance in probabilistic forecasting. Results confirm that ensembling consistently improves prediction—especially probabilistic calibration—but high accuracy typically incurs high cost. Our framework delivers a scalable, cost-effective ensemble strategy for resource-constrained deployment.
This paper addresses insufficient risk diversification in portfolio optimization by proposing a structured ensemble learning framework based on multi-hypothesis prediction, unifying asset selection and weight optimization within a prediction-to-optimization pipeline subject to diversity constraints. Its key contributions include: (i) explicitly linking ensemble loss decomposition theory to portfolio diversification; (ii) introducing a pre-screening mechanism that dynamically balances predictive accuracy against structural diversity; and (iii) constructing a parameterized prediction set with controllable diversity and a supervised ensemble combiner (e.g., equal-weighted aggregation under squared loss). Empirical evaluation across over two decades of S&P 500 constituents and a global bond dataset comprising 1,300 instruments demonstrates significantly expanded achievable diversification bounds. The framework delivers robust, state-of-the-art performance in both single-period and multi-period portfolio allocation tasks.
This paper addresses the challenge of quantifying individual model contributions in multi-model ensemble forecasting. We propose an interpretable attribution framework grounded in Shapley values from cooperative game theory—the first application of Shapley values to ensemble importance assessment. To ensure scalability and theoretical rigor, we introduce two efficient algorithms: Leave-One-Model-Out (LOMO) and Leave-All-Subsets-of-Models-Out (LASMO). By integrating error similarity analysis and Monte Carlo approximation, we significantly reduce computational complexity. Evaluated on the US COVID-19 mortality prediction task, our method identifies models with low standalone accuracy but high collaborative value—revealing complementary and redundant interactions among models that conventional accuracy metrics fail to capture. The framework advances ensemble interpretability and informs principled model selection, establishing a new paradigm for explainable ensemble learning.
Diffusion models for weather forecasting suffer from autoregressive rollout, high computational cost, and error accumulation at high temporal resolution. To address these issues, this paper proposes Continuous Ensemble Diffusion (CED), a parallel ensemble forecasting framework based on continuous-time modeling. CED treats meteorological fields as a continuous-time stochastic process governed by an implicit ordinary differential equation (ODE), enabling parallel generation of temporally coherent ensemble trajectories in a single forward pass—eliminating autoregressive iteration entirely. The method integrates implicit ODE solvers, ensemble probabilistic calibration, and parallel denoising, supporting arbitrary temporal resolution and flexible hybridization with conventional rollout. On global weather forecasting benchmarks, CED achieves state-of-the-art performance: it improves skill scores significantly, reduces the Continuous Ranked Probability Score (CRPS) by 12.3%, and exhibits superior spread-error consistency.
This paper addresses the Stochastic Black-Box Optimization Selection (SBOS) problem: identifying, among multiple stochastic systems with continuous decision variables, the system whose optimal decision yields the best expected performance—without prior knowledge and under a finite sampling budget. We propose the first formal SBOS framework that jointly integrates intra-system optimization via stochastic gradient descent and inter-system comparison via sequential elimination, enabling synergistic optimization across both levels. We theoretically establish that the probability of incorrect selection converges exponentially with the sampling budget. Empirical evaluation across three real-world SBOS scenarios demonstrates that our method significantly reduces the misselection probability and maintains robust superiority across varying budget sizes and problem dimensions.
Quantifying the contribution of individual submodels to the overall predictive performance of ensemble systems is crucial for enhancing interpretability and construction efficiency. This work proposes and implements a unified R package that, for the first time in the R ecosystem, systematically supports model importance assessment across diverse ensemble methods under both point and probabilistic forecasting frameworks, with full compatibility with the hubverse infrastructure. The package offers flexible importance metrics and robust handling of missing values, substantially improving the understanding of submodel roles. It thereby empowers researchers to efficiently construct, diagnose, and optimize ensemble forecasting systems.
This study addresses the common issue in ensemble forecasting wherein insufficiently rapid growth of ensemble spread leads to inadequate representation of uncertainty. Using the Lorenz '96 system, the work systematically disentangles intrinsic variability, initial condition perturbations, and stochastic model uncertainty to evaluate how various ensemble configurations and parameterization schemes influence spread evolution. It introduces novel Bayesian and streaming stochastic parameterizations featuring temporally coherent structures, revealing that perturbations primarily govern the rate of trajectory decorrelation rather than long-term variance. The analysis further elucidates the interaction mechanisms among distinct uncertainty sources. Experimental results demonstrate that the proposed methods significantly enhance early spread growth and improve consistency between ensemble spread and forecast error, thereby offering theoretical insights and practical guidance for uncertainty modeling in numerical weather prediction systems.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.