confidence-weighted ensembling

Designs and implements methods that combine multiple predictive models' outputs into a single prediction by weighting each model's contribution on a per-sample basis according to an estimate of that model's confidence, plausibility, or evidence (e.g., softmax averaging, reservoir ensemble averaging, plausibility-weighted ensemble predictions). Builds algorithms to compute or calibrate per-sample confidence scores, to fuse adapted and base model outputs, to downweight overconfident predictions, and to analyze the ensemble's behavior and robustness under distribution shift.

confidence-weightedensembling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-Hypothesis Prediction for Portfolio Optimization: A Structured Ensemble Learning Approach to Risk Diversification

Jan 07, 2025
AR
Alejandro Rodriguez Dominguez
🏛️ University of Reading | Miralta Finance Bank S.A.

This paper addresses insufficient risk diversification in portfolio optimization by proposing a structured ensemble learning framework based on multi-hypothesis prediction, unifying asset selection and weight optimization within a prediction-to-optimization pipeline subject to diversity constraints. Its key contributions include: (i) explicitly linking ensemble loss decomposition theory to portfolio diversification; (ii) introducing a pre-screening mechanism that dynamically balances predictive accuracy against structural diversity; and (iii) constructing a parameterized prediction set with controllable diversity and a supervised ensemble combiner (e.g., equal-weighted aggregation under squared loss). Empirical evaluation across over two decades of S&P 500 constituents and a global bond dataset comprising 1,300 instruments demonstrates significantly expanded achievable diversification bounds. The framework delivers robust, state-of-the-art performance in both single-period and multi-period portfolio allocation tasks.

Introduces asset selection prioritizing diversity over predicted returnsLinks predictors' diversity to risk diversification via ensemble learningProposes a unified framework for portfolio allocation and optimization

Random Subset Averaging

Dec 27, 2025
WC
Wenhao Cui
🏛️ Beihang University | Tsinghua University

To address prediction instability and inaccuracy under high-dimensional, strongly correlated covariates, this paper proposes the Random Subset Averaging (RSA) ensemble method: base models are constructed via binomial random sampling, and adaptively fused through two rounds of data-driven weighting—structurally analogous to a two-layer neural network. Theoretical contributions include: (i) the first asymptotic optimality guarantee for two-stage adaptive weighting; (ii) allowance for fully data-dependent weights in the first stage; and (iii) tighter finite-sample risk bounds under orthogonal design. RSA integrates cross-validation-based hyperparameter tuning with dual-stage weighting, ensuring both computational feasibility and statistical rigor. Empirical results demonstrate that RSA significantly outperforms LASSO, random forests, and state-of-the-art ensemble methods across diverse sparse and highly correlated settings. Its practical robustness and superiority are further validated in financial return forecasting.

Develops ensemble method for high-dimensional correlated dataOptimizes predictions via two-round weighting and cross-validationOutperforms existing methods in simulations and financial forecasting

On Arbitrary Predictions from Equally Valid Models

Jul 25, 2025
SL
Sarah Lockfisch
🏛️ Technical University of Munich | TUM University Hospital | Imperial College London | Munich Center for Machine Learning | Google DeepMind

Medical diagnostics face the “model multiplicity” problem: multiple machine learning models with comparable aggregate performance may yield contradictory predictions for the same patient; standard evaluation metrics fail to identify the optimal model, and individual predictions are susceptible to arbitrary choices made during development—undermining diagnostic reliability. This paper proposes a mitigation framework based on small-scale ensembling and predictive abstention, which leverages cross-model consensus analysis to identify high-agreement predictions for automated diagnosis while explicitly deferring ambiguous cases to clinical experts. Its key innovation lies in explicitly modeling model uncertainty as an actionable clinical decision signal. Experiments demonstrate that the approach significantly reduces prediction multiplicity and improves diagnostic consistency, without substantially increasing model complexity.

Analyzing predictive multiplicity in medical machine learning modelsIdentifying arbitrary predictions from equally valid model choicesProposing ensemble strategies to mitigate diagnostic unreliability

Rethinking Weight-Averaged Model-merging

Nov 14, 2024
HW
Hu Wang
🏛️ Mohamed bin Zayed University of Artificial Intelligence | University of Surrey

This paper addresses the opacity of weight averaging—a widely used model merging technique—by systematically investigating its effectiveness from three perspectives: (1) revealing that model weights inherently encode structured semantic patterns, enabling interpretability; (2) theoretically and empirically contrasting weight averaging with feature averaging, establishing its implicit regularization property; and (3) demonstrating strong prediction stability under parameter-scale variations, indicating scale robustness. Through weight visualization, mathematical modeling, and extensive cross-architecture and cross-dataset experiments, we provide the first mechanistic decomposition of this “black-box” operation. Our work delivers rigorous interpretability evidence and practical guidelines for training-free model fusion. All code is publicly released, advancing both theoretical understanding and engineering deployment of parameter-space ensembling.

Compares weight averaging versus feature averaging in model ensembles.Explores mechanisms behind weight-averaged model-merging effectiveness.Investigates parameter magnitude impact on model-merging prediction stability.

Latest Papers

What's happening recently
View more

This work addresses the challenge of ensuring reliable calibration in Mixture-of-Experts (MoE) models under distribution shift, where well-calibrated individual experts do not guarantee a well-calibrated aggregate prediction—particularly under soft routing. The study systematically analyzes the interplay between hard and soft routing mechanisms and expert calibration, revealing fundamental differences in their impact on overall model calibration. To mitigate this issue, the authors propose a novel adversarial reweighting strategy applied to the routed aggregation output, which significantly enhances calibration robustness without compromising accuracy. Extensive experiments demonstrate that the method consistently improves the trade-off between accuracy and calibration across diverse architectures, tasks, and distribution shift scenarios, with especially pronounced gains on challenging data subsets.

calibrationdistribution shiftmixture-of-experts

This work proposes an interpretable statistical inference framework that decomposes predictive scoring functions into three components—calibration error, discrimination ability, and uncertainty—applicable to multi-step-ahead point forecasts such as means and quantiles, and compatible with both smooth and nonsmooth scoring rules. Building on linear recalibration and integrating Mincer–Zarnowitz regression with asymptotic inference theory, the method delivers the first fully interpretable tripartite decomposition for general scoring functions, unifying and extending classical calibration tests and predictive performance evaluation. Empirical applications to inflation surveys and financial risk models reveal critical discrepancies obscured by aggregate scores, exposing a misalignment between backtesting practices and predictive accuracy in banking regulation, thereby substantially enhancing the informativeness and statistical power of forecast evaluation.

discriminationforecast calibrationpredictive assessment

This work proposes an efficient Bayesian deep ensemble method to address the limitations of existing deep ensembles in uncertainty calibration and interpretability. By leveraging low-dimensional predictive representations, independent training strategies, and a closed-form Bayesian linear regression aggregation mechanism, the approach enables analytical posterior weight inference while maintaining high predictive accuracy. This significantly enhances model interpretability and yields well-calibrated uncertainty estimates. Notably, the computational complexity shifts from scaling with dataset size to depending only on ensemble size, substantially improving scalability. Empirical evaluations on standard regression benchmarks demonstrate that the method achieves state-of-the-art predictive performance while providing reliable and properly calibrated uncertainty quantification.

Bayesian deep ensemblescomputational efficiencyinterpretability

This work addresses the limited predictive diversity and poor out-of-distribution (OOD) uncertainty quantification of single pretrained models under distribution shift. The authors propose a Perturb-and-Correct approach that constructs a posterior ensemble using only a single model by applying random perturbations to hidden layers of the pretrained network and subsequently correcting them via least-squares affine transformations. This method uniquely exploits the affine redundancy inherent in neural networks, enhancing OOD prediction diversity and uncertainty calibration without compromising in-distribution performance. Empirical evaluations demonstrate that the proposed technique achieves a superior or competitive trade-off between in-distribution accuracy and OOD detection on MuJoCo dynamics prediction and CIFAR-10 OOD benchmarks compared to existing posterior ensemble baselines.

distribution shiftepistemic diversitymodel calibration

Hot Scholars

FL

Fang Liu

Professor of Xidian University
AIMachine LearningPattern RecognitionImage Processing
LL

Lewei Lu

Research Director (We're Hiring, luotto@sensetime.com) @ SenseTime Research
Computer VisionDeep Learning
AB

Atif Belal

Student, Ecole de Technologie Superieur
Machine LearningDeep LearningComputer Vision
EG

Eric Granger

Professor of Systems Engineering, École de technologie supérieure, LIVIA, ILLS, REPARTI
Machine LearningComputer VisionPattern RecognitionAffective Computing
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing