Score
Designs and implements methods that combine multiple predictive models' outputs into a single prediction by weighting each model's contribution on a per-sample basis according to an estimate of that model's confidence, plausibility, or evidence (e.g., softmax averaging, reservoir ensemble averaging, plausibility-weighted ensemble predictions). Builds algorithms to compute or calibrate per-sample confidence scores, to fuse adapted and base model outputs, to downweight overconfident predictions, and to analyze the ensemble's behavior and robustness under distribution shift.
This paper addresses insufficient risk diversification in portfolio optimization by proposing a structured ensemble learning framework based on multi-hypothesis prediction, unifying asset selection and weight optimization within a prediction-to-optimization pipeline subject to diversity constraints. Its key contributions include: (i) explicitly linking ensemble loss decomposition theory to portfolio diversification; (ii) introducing a pre-screening mechanism that dynamically balances predictive accuracy against structural diversity; and (iii) constructing a parameterized prediction set with controllable diversity and a supervised ensemble combiner (e.g., equal-weighted aggregation under squared loss). Empirical evaluation across over two decades of S&P 500 constituents and a global bond dataset comprising 1,300 instruments demonstrates significantly expanded achievable diversification bounds. The framework delivers robust, state-of-the-art performance in both single-period and multi-period portfolio allocation tasks.
To address prediction instability and inaccuracy under high-dimensional, strongly correlated covariates, this paper proposes the Random Subset Averaging (RSA) ensemble method: base models are constructed via binomial random sampling, and adaptively fused through two rounds of data-driven weighting—structurally analogous to a two-layer neural network. Theoretical contributions include: (i) the first asymptotic optimality guarantee for two-stage adaptive weighting; (ii) allowance for fully data-dependent weights in the first stage; and (iii) tighter finite-sample risk bounds under orthogonal design. RSA integrates cross-validation-based hyperparameter tuning with dual-stage weighting, ensuring both computational feasibility and statistical rigor. Empirical results demonstrate that RSA significantly outperforms LASSO, random forests, and state-of-the-art ensemble methods across diverse sparse and highly correlated settings. Its practical robustness and superiority are further validated in financial return forecasting.
Medical diagnostics face the “model multiplicity” problem: multiple machine learning models with comparable aggregate performance may yield contradictory predictions for the same patient; standard evaluation metrics fail to identify the optimal model, and individual predictions are susceptible to arbitrary choices made during development—undermining diagnostic reliability. This paper proposes a mitigation framework based on small-scale ensembling and predictive abstention, which leverages cross-model consensus analysis to identify high-agreement predictions for automated diagnosis while explicitly deferring ambiguous cases to clinical experts. Its key innovation lies in explicitly modeling model uncertainty as an actionable clinical decision signal. Experiments demonstrate that the approach significantly reduces prediction multiplicity and improves diagnostic consistency, without substantially increasing model complexity.
This paper addresses the opacity of weight averaging—a widely used model merging technique—by systematically investigating its effectiveness from three perspectives: (1) revealing that model weights inherently encode structured semantic patterns, enabling interpretability; (2) theoretically and empirically contrasting weight averaging with feature averaging, establishing its implicit regularization property; and (3) demonstrating strong prediction stability under parameter-scale variations, indicating scale robustness. Through weight visualization, mathematical modeling, and extensive cross-architecture and cross-dataset experiments, we provide the first mechanistic decomposition of this “black-box” operation. Our work delivers rigorous interpretability evidence and practical guidelines for training-free model fusion. All code is publicly released, advancing both theoretical understanding and engineering deployment of parameter-space ensembling.
This work addresses the challenge of ensuring reliable calibration in Mixture-of-Experts (MoE) models under distribution shift, where well-calibrated individual experts do not guarantee a well-calibrated aggregate prediction—particularly under soft routing. The study systematically analyzes the interplay between hard and soft routing mechanisms and expert calibration, revealing fundamental differences in their impact on overall model calibration. To mitigate this issue, the authors propose a novel adversarial reweighting strategy applied to the routed aggregation output, which significantly enhances calibration robustness without compromising accuracy. Extensive experiments demonstrate that the method consistently improves the trade-off between accuracy and calibration across diverse architectures, tasks, and distribution shift scenarios, with especially pronounced gains on challenging data subsets.
This work proposes an interpretable statistical inference framework that decomposes predictive scoring functions into three components—calibration error, discrimination ability, and uncertainty—applicable to multi-step-ahead point forecasts such as means and quantiles, and compatible with both smooth and nonsmooth scoring rules. Building on linear recalibration and integrating Mincer–Zarnowitz regression with asymptotic inference theory, the method delivers the first fully interpretable tripartite decomposition for general scoring functions, unifying and extending classical calibration tests and predictive performance evaluation. Empirical applications to inflation surveys and financial risk models reveal critical discrepancies obscured by aggregate scores, exposing a misalignment between backtesting practices and predictive accuracy in banking regulation, thereby substantially enhancing the informativeness and statistical power of forecast evaluation.
This work proposes an efficient Bayesian deep ensemble method to address the limitations of existing deep ensembles in uncertainty calibration and interpretability. By leveraging low-dimensional predictive representations, independent training strategies, and a closed-form Bayesian linear regression aggregation mechanism, the approach enables analytical posterior weight inference while maintaining high predictive accuracy. This significantly enhances model interpretability and yields well-calibrated uncertainty estimates. Notably, the computational complexity shifts from scaling with dataset size to depending only on ensemble size, substantially improving scalability. Empirical evaluations on standard regression benchmarks demonstrate that the method achieves state-of-the-art predictive performance while providing reliable and properly calibrated uncertainty quantification.
This work addresses the limited predictive diversity and poor out-of-distribution (OOD) uncertainty quantification of single pretrained models under distribution shift. The authors propose a Perturb-and-Correct approach that constructs a posterior ensemble using only a single model by applying random perturbations to hidden layers of the pretrained network and subsequently correcting them via least-squares affine transformations. This method uniquely exploits the affine redundancy inherent in neural networks, enhancing OOD prediction diversity and uncertainty calibration without compromising in-distribution performance. Empirical evaluations demonstrate that the proposed technique achieves a superior or competitive trade-off between in-distribution accuracy and OOD detection on MuJoCo dynamics prediction and CIFAR-10 OOD benchmarks compared to existing posterior ensemble baselines.