train boosting meta-learners

Designs and trains stacked ensemble models that take base-model outputs (predictions, probabilities, or uncertainty estimates — including from test‑time augmentation) as features and fit gradient‑boosted decision tree meta‑learners (level‑1 or level‑2) to produce final predictions. Builds and validates meta‑classifiers/blenders for calibration and robustness using out‑of‑fold training, cross‑validation, and leakage‑mitigation practices.

trainboostingmeta-learners

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Dynamic Meta-Learning for Adaptive XGBoost-Neural Ensembles

Sep 30, 2025
AS
Arthur Sedek
🏛️ IMDEX Limited

This work addresses the rigidity in model selection and poor interpretability inherent in conventional XGBoost–neural network ensembles. We propose an adaptive fusion framework grounded in dynamic meta-learning, which jointly leverages uncertainty quantification and feature importance as dual control signals to guide fine-grained scheduling and weighted integration of the two base models at inference time. Our key innovation lies in co-modeling uncertainty estimates and interpretability-aware metrics—specifically, feature importance—within the meta-learner’s decision process, thereby simultaneously enhancing predictive performance and decision transparency. Extensive experiments across multiple benchmark datasets demonstrate that our method consistently outperforms static ensembles and individual baselines, achieving average accuracy gains of 2.1–4.7 percentage points. Moreover, it provides auditable, instance-level rationale for model selection and feature-level attribution, supporting both reliability assessment and human-understandable explanations.

Adaptive model selection using uncertainty quantification techniquesDynamic ensemble framework combining XGBoost and neural networksEnhanced predictive performance and interpretability across diverse datasets

Deep ensemble methods often improve predictive performance, yet they suffer from three practical limitations: redundancy among base models that inflates computational cost and degrades conditioning, unstable weighting under multicollinearity, and overfitting in meta-learning pipelines. We propose a regularized meta-learning framework that addresses these challenges through a four-stage pipeline combining redundancy-aware projection, statistical meta-feature augmentation, and cross-validated regularized meta-models (Ridge, Lasso, and ElasticNet). Our multi-metric de-duplication strategy removes near-collinear predictors using correlation and MSE thresholds ($\tau_{\text{corr}}=0.95$), reducing the effective condition number of the meta-design matrix while preserving predictive diversity. Engineered ensemble statistics and interaction terms recover higher-order structure unavailable to raw prediction columns. A final inverse-RMSE blending stage mitigates regularizer-selection variance. On the Playground Series S6E1 benchmark (100K samples, 72 base models), the proposed framework achieves an out-of-fold RMSE of 8.582, improving over simple averaging (8.894) and conventional Ridge stacking (8.627), while matching greedy hill climbing (8.603) with substantially lower runtime (4 times faster). Conditioning analysis shows a 53.7\% reduction in effective matrix condition number after redundancy projection. Comprehensive ablations demonstrate consistent contributions from de-duplication, statistical meta-features, and meta-ensemble blending. These results position regularized meta-learning as a stable and deployment-efficient stacking strategy for high-dimensional ensemble systems.

deep ensemblemeta-learningmulticollinearity

Stacked conformal prediction

May 18, 2025
PC
Paulo C. Marques
🏛️ Insper Institute of Education and Research

This work addresses the low data efficiency and high computational overhead of conformal prediction caused by its reliance on an independent calibration set. We propose an end-to-end stacked ensemble conformalization method that embeds conformal prediction into a stacking framework, employing a lightweight top-layer meta-learner to directly model nonconformity scores—eliminating the need for calibration-set splitting while achieving approximate marginal coverage guarantees. Theoretical analysis establishes statistical validity under mild assumptions. Empirical evaluation across multiple benchmark datasets demonstrates that our approach yields more stable coverage, reduced calibration error, and significantly lower inference and conformalization costs compared to standard inductive conformal prediction. To the best of our knowledge, this is the first method to jointly realize end-to-end conformalization of stacked ensembles, effectively balancing statistical rigor with engineering practicality.

Achieving marginal validity without calibration samplesComparing favorably to standard inductive methodsConformalizing stacked ensemble models efficiently

An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient-Boosting and Fuzzy Rule-Based Models

Nov 01, 2024
JL
Jinbo Li
🏛️ China Unicom | Henan University | University of Macau | University of Alberta | Nantong University

Fuzzy rule models offer strong interpretability but suffer from poor scalability and susceptibility to overfitting in complex tasks and large-scale data scenarios. To address these limitations, this paper proposes a novel ensemble framework integrating gradient boosting with fuzzy rule-based base learners. We introduce a dynamic control factor that adaptively adjusts the weights of fuzzy base models in each boosting iteration, simultaneously serving as a regularizer and performance optimizer. Additionally, we design a validation-set-driven, sample-level correction mechanism to enhance generalization and ensemble diversity. Experimental results demonstrate that our approach significantly mitigates overfitting, reduces rule complexity (e.g., fewer rules and shorter antecedents), and preserves high model interpretability and maintainability. The method thus provides a practical pathway for deploying interpretable AI in complex industrial applications.

Addresses overfitting and complexity in fuzzy rule-based systems using dynamic controlCombines gradient boosting with fuzzy models to enhance performance and interpretabilityImproves scalability and maintenance of ensemble models through adaptive tuning mechanisms

This paper addresses the lack of theoretical guarantees for stacking generalization in temporal probabilistic forecasting. We propose a quantile–time–item adaptive weighted ensemble method. First, we establish a tight generalization error bound for cross-validation-driven stacked ensembles—improving upon Van der Laan et al. (2007)—via empirical process theory and concentration inequalities. Second, our method introduces a structured weight modeling family that enables quantile-aware dynamic weight learning, ensuring both theoretical rigor and practical flexibility. Third, extensive experiments on multiple probabilistic forecasting benchmarks demonstrate significant improvements over conventional ensemble baselines. Empirical results confirm that dynamic weight adaptation is critical for enhancing both predictive accuracy and calibration.

Extending and strengthening existing theoretical results on ensemble performanceProposing a novel ensemble method for probabilistic time series forecastingProving theoretical guarantees for stacked generalization in ensemble learning

Latest Papers

What's happening recently
View more

This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.

binary classificationcomputational costearly stopping

This work addresses the challenges of multiclass classification—such as high inter-class similarity, class imbalance, and significant distributional discrepancies—where a single model often struggles to simultaneously capture complex functional relationships and sharp decision boundaries. The authors propose LFS-FRAME, a novel framework featuring a leakage-free stacking mechanism that synergistically combines the global function approximation capability of Kolmogorov–Arnold Networks (KANs) with the local rule-learning strength of XGBoost. By employing rigorous out-of-fold validation to generate unbiased meta-features, the framework effectively integrates functional learning with probabilistic meta-learning. Evaluated across multiple datasets, the method achieves 89.85% accuracy in main-group identification and 81.74% in sub-group recognition, substantially outperforming strong single-model baselines and enhancing both robustness and generalization in multiclass classification tasks.

class imbalancedata distribution variabilitygeneralization

This study addresses the limitations of traditional decision trees lacking meta-learning priors and the difficulty of foundation models in generating auditable, standalone models. To this end, we propose MotherTree, which employs a tabular Transformer architecture pretrained on synthetic data to directly output hard axis-aligned decision trees in a single forward pass. Furthermore, it introduces a novel meta-learning paradigm that eliminates the need for reference tree supervision, effectively transforming black-box predictions into interpretable, independently deployable traditional decision trees. Experimental results demonstrate that MotherTree achieves performance comparable to classical algorithms on standard benchmarks while significantly outperforming gradient-based learning from scratch. Additionally, it serves as a strong initializer that enhances fine-tuning efficacy, successfully balancing model performance with transparency.

decision treemeta-learningsmall-sample regimes

This study addresses the insufficient robustness of ensemble models against component failures and the evaluation bias arising from single-corpus assessments in AI-generated image detection. Through rigorous controlled experiments and preprocessing path analysis, it reveals the limitations of stacking ensembles that rely on in-domain retraining. Revising earlier conclusions, this work proposes an architecture-diversity-based majority voting rule, validated via a gradient-boosted meta-learner and McNemar’s test. It demonstrates that uncalibrated ensembles underperform optimal individual models, thereby establishing the necessity of domain adaptation. Furthermore, the proposed voting rule effectively prevents cascading failures without retraining while exposing deficiencies in abstention handling and potential JPEG compression biases. All code and corrected results have been made publicly available.

AI-generated image detectioncomponent failuredetector abstention

Hot Scholars

ZM

Zeyuan Ma

South China University of Technology
Meta-Black-Box OptimizationReinforcement LearningLearning to Optimize
BK

Boris Knyazev

Research Scientist, Samsung - SAIT AI Lab
Machine LearningComputer VisionArtificial Intelligence
EB

Eugene Belilovsky

Associate Professor, Concordia University and Mila Quebec AI Institute
Distributed LearningContinual LearningFederated LearningLearned Optimizers
JC

Jesse C. Cresswell

Layer 6 AI
Trustworthy MLDeep Generative ModellingQuantum Information
CG

Carlos Guestrin

Professor, Stanford University
Machine LearningDistributed SystemsArtificial IntelligenceParallel Algorithms