model ensembling

Design and build systems that combine multiple trained models or their embeddings into a single predictive or representational output, implementing methods for selecting, weighting, training, and compressing ensemble members. This includes creating heterogeneous ensembles across architectures or pretrained backbones, seed- or pretraining-diverse ensembles, embedding averaging or fusion, per-class weighted or voting/consensus aggregation (including committees of large language models), and evaluating ensemble selection, weighting, and regularization effects.

modelensembling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Dynamic Post-Hoc Neural Ensemblers

Oct 06, 2024
SP
Sebastian Pineda Arango
🏛️ University of Freiburg | ELLIS Institute Tübingen | University of Technology Nürnberg

Existing ensemble methods (e.g., greedy or random ensembling) employ static, sample-agnostic base-model weights, limiting expressive capacity and generalization. This paper proposes the Dynamic Posterior Neural Ensemble (DPNE), the first framework to introduce Random Prediction Dropout (RPD)—a theoretically grounded diversity regularization mechanism that provably enhances base-model diversity under a derived lower bound. DPNE employs a neural architecture to learn sample-wise adaptive weights end-to-end, eliminating restrictive pre-specified weight structures. Extensive experiments across CV, NLP, and tabular tasks demonstrate that DPNE significantly outperforms state-of-the-art ensemble baselines, effectively mitigating overfitting while improving both in-distribution accuracy and out-of-distribution robustness. The core contributions are: (1) dynamic posterior weight modeling conditioned on input samples, and (2) theoretically guaranteed diversity regularization via RPD.

Addressing limitations of constant-weight ensembling approachesEnhancing accuracy and robustness of ensemble methodsImproving generalization via regularized dynamic ensembling

Divide, Specialize, and Route: A New Approach to Efficient Ensemble Learning

Jun 25, 2025
JP
Jakub Piwko
🏛️ Warsaw University of Technology

Traditional ensemble methods (e.g., Bagging, Boosting) suffer from high computational overhead and poor adaptability to heterogeneous data distributions. To address these limitations, we propose Hellsemble—a computationally efficient, progressive ensemble framework for binary classification. Hellsemble first partitions instances into dynamic “difficulty tiers” based on instance-level difficulty estimation; a learned router then progressively routes challenging samples across tiers to specialized base learners, enabling model specialization and load balancing. Its core innovations are a difficulty-driven dynamic routing mechanism and a progressive training paradigm, jointly enhancing interpretability, generalization, and inference efficiency. Evaluated on the OpenML-CC18 and Tabzilla benchmarks, Hellsemble consistently outperforms state-of-the-art ensemble methods, achieving an average accuracy gain of 2.1% and accelerating inference by 3.4×.

Address high computational cost in ensemble learningEnhance interpretability in binary classification frameworksImprove adaptability to heterogeneous data distributions

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

Mar 13, 2025
MA
Mari Ashiga
🏛️ University of West London | Turing Intelligence Technology Limited | University of Leeds

This paper addresses three critical challenges in large language models (LLMs): output inconsistency across single-model generations, limited pattern diversity due to inherent linguistic biases, and data privacy risks plus industrial integration barriers stemming from closed-source architectures. We systematically survey and reconceptualize LLM ensemble learning methodologies, categorizing existing techniques into seven classes and identifying four high-performance paradigms—weight merging, knowledge fusion, mixture-of-experts, and reward-based ensembling. We further propose a cross-modal transfer pathway to extend ensemble models to multimodal settings. Through joint analysis of modeling pipelines, training strategies, and output characteristics, empirical evaluation demonstrates that ensemble methods significantly improve generation diversity, quality, and task-adaptive flexibility. Our work provides both theoretical foundations and practical guidelines for industrial-scale LLM selection, customization, and deployment.

Address inconsistencies in text and code generation by LLMs.Explore ensemble methods to enhance LLM flexibility and quality.Overcome biases and limited diversity in LLM outputs.

Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling

Oct 03, 2024
YY
Yuxuan Yao
🏛️ City University of Hong Kong | Huawei | North China University of Technology

Existing LLM ensemble methods often overlook model compatibility and rely on full-vocabulary probability alignment, resulting in low efficiency and high computational overhead. This work identifies model compatibility as a critical determinant of ensemble performance and proposes Union Top-k Ensemble—a novel paradigm that first selects a compatible subset of models and then performs lightweight probability aggregation solely over the union of each model’s top-k predicted tokens, thereby avoiding full-vocabulary alignment. Our approach establishes the first token-level, compatibility-aware ensemble framework that requires no retraining. Evaluated across multiple benchmark tasks, it significantly outperforms state-of-the-art ensemble methods in both accuracy and robustness while reducing computational cost by 30%–50%.

Develop efficient model selection strategyIdentify model compatibility for effective ensemblingIntroduce UniTE for top-k token union ensembling

The cost of ensembling: is it always worth combining?

Jun 05, 2025
MZ
Marco Zanotti
🏛️ University of Milano-Bicocca

Time-series ensemble forecasting faces a critical trade-off between predictive accuracy and computational cost. This paper systematically evaluates ten base models and eight ensemble strategies on the M5 and VN1 retail datasets, measuring performance in point forecasting (RMSE) and probabilistic forecasting (CRPS), alongside computational overhead. Methodologically, we analyze ensemble size scalability, propose an “efficiency-driven ensemble” paradigm, and assess downsampling-based retraining frequency reduction. Key contributions: (1) Ensembles of only two to three models achieve near-optimal accuracy; (2) The efficiency-driven paradigm reduces average computational cost by over 40% while retaining ≥95% of baseline accuracy; (3) Reducing retraining frequency cuts training overhead by up to 70%, with negligible impact on point forecasts and robust performance in probabilistic forecasting. Results confirm that ensembling consistently improves prediction—especially probabilistic calibration—but high accuracy typically incurs high cost. Our framework delivers a scalable, cost-effective ensemble strategy for resource-constrained deployment.

Assessing impact of retraining frequency on cost and accuracy in forecastingEvaluating trade-offs between forecast accuracy and computational cost in ensemble learningIdentifying optimal small ensembles for cost-efficient time series forecasting

Latest Papers

What's happening recently
View more

This study investigates optimal parallel heterogeneous ensemble strategies for small- to medium-scale tabular classification tasks. Through systematic experiments on 56 OpenML CC18 datasets, the authors find that the instability of Blending and Stacking is mutually independent, and that Robust Soft Voting significantly outperforms Hard Voting in multi-class settings. Building on these insights, they propose a robust ensemble strategy and validate its effectiveness on the TabArena platform. The recommended approach significantly surpasses the Single Best method across 28 additional tasks and matches or exceeds the performance of all individual ensemble techniques considered.

classification performanceensemble selectionheterogeneous ensemble

This work addresses the lack of systematic understanding in efficiently fusing large language models (LLMs) fine-tuned with lightweight adapters in multi-task learning, particularly regarding the trade-offs among ensembling, merging, and routing strategies. The study systematically evaluates three parameter-efficient fusion approaches—output ensembling, parameter averaging, and input-dependent routing—and demonstrates that non-uniform fusion consistently outperforms uniform methods, with routing yielding significant performance gains despite its higher computational cost. To reconcile this efficiency–performance trade-off, the authors propose a low-overhead expert selection mechanism that combines clustering with greedy subset selection, achieving near-optimal performance while substantially reducing computational overhead, thereby striking an effective balance between model efficacy and efficiency.

ensemblingmergingmulti-task learning

This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.

binary classificationcomputational costearly stopping

This work addresses the challenges of high computational costs and inefficient subset selection in large-scale image training. It proposes the SCOre-Stratified Selection (SCOSS) framework, which constructs a coreset through score-stratified sampling and integrates predictions from models trained on multiple independent subsets. By innovatively combining stratified sampling with ensemble learning, SCOSS significantly enhances the stability and generalization capability of coreset selection. Experimental results demonstrate that, across various sampling ratios, SCOSS-based coresets achieve state-of-the-art performance on the Simple Graph Convolution (SGC) model, surpassing support vector machines (SVMs) with only a small number of labeled samples while striking an excellent balance between accuracy and efficiency.

coreset selectionimage classificationlarge-scale datasets

While ensemble models achieve strong performance, their large scale poses significant challenges in deployment, interpretability, and robustness verification. This work proposes PACE, a novel framework that uniquely integrates model compression and pruning through a two-stage alternating strategy. First, it enhances ensemble diversity by generating a set of diverse weak learners guided by a theoretically grounded active learning mechanism. Subsequently, it applies fidelity-aware pruning to the expanded ensemble, enabling controlled preservation of the original model’s behavioral characteristics. Extensive experiments demonstrate that PACE consistently outperforms existing compression and pruning methods across multiple tasks, achieving higher compression ratios while more faithfully retaining the behavior of the original ensemble.

compressiondeploymentensemble models

Hot Scholars

RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
XW

Xinrun Wang

Singapore Management University
Reinforcement LearningMulti-Agent SystemsDistributed Foundation Agents
JK

Jihie Kim

Dongguk University
Artificial IntelligenceComputer EducationHuman Computer InteractionNLP
SH

Shuyue Hu

Shanghai Artificial Intelligence Lab
multiagent systemlarge language modelgame theory