Score
Design and build systems that combine multiple trained models or their embeddings into a single predictive or representational output, implementing methods for selecting, weighting, training, and compressing ensemble members. This includes creating heterogeneous ensembles across architectures or pretrained backbones, seed- or pretraining-diverse ensembles, embedding averaging or fusion, per-class weighted or voting/consensus aggregation (including committees of large language models), and evaluating ensemble selection, weighting, and regularization effects.
Existing ensemble methods (e.g., greedy or random ensembling) employ static, sample-agnostic base-model weights, limiting expressive capacity and generalization. This paper proposes the Dynamic Posterior Neural Ensemble (DPNE), the first framework to introduce Random Prediction Dropout (RPD)—a theoretically grounded diversity regularization mechanism that provably enhances base-model diversity under a derived lower bound. DPNE employs a neural architecture to learn sample-wise adaptive weights end-to-end, eliminating restrictive pre-specified weight structures. Extensive experiments across CV, NLP, and tabular tasks demonstrate that DPNE significantly outperforms state-of-the-art ensemble baselines, effectively mitigating overfitting while improving both in-distribution accuracy and out-of-distribution robustness. The core contributions are: (1) dynamic posterior weight modeling conditioned on input samples, and (2) theoretically guaranteed diversity regularization via RPD.
Traditional ensemble methods (e.g., Bagging, Boosting) suffer from high computational overhead and poor adaptability to heterogeneous data distributions. To address these limitations, we propose Hellsemble—a computationally efficient, progressive ensemble framework for binary classification. Hellsemble first partitions instances into dynamic “difficulty tiers” based on instance-level difficulty estimation; a learned router then progressively routes challenging samples across tiers to specialized base learners, enabling model specialization and load balancing. Its core innovations are a difficulty-driven dynamic routing mechanism and a progressive training paradigm, jointly enhancing interpretability, generalization, and inference efficiency. Evaluated on the OpenML-CC18 and Tabzilla benchmarks, Hellsemble consistently outperforms state-of-the-art ensemble methods, achieving an average accuracy gain of 2.1% and accelerating inference by 3.4×.
This paper addresses three critical challenges in large language models (LLMs): output inconsistency across single-model generations, limited pattern diversity due to inherent linguistic biases, and data privacy risks plus industrial integration barriers stemming from closed-source architectures. We systematically survey and reconceptualize LLM ensemble learning methodologies, categorizing existing techniques into seven classes and identifying four high-performance paradigms—weight merging, knowledge fusion, mixture-of-experts, and reward-based ensembling. We further propose a cross-modal transfer pathway to extend ensemble models to multimodal settings. Through joint analysis of modeling pipelines, training strategies, and output characteristics, empirical evaluation demonstrates that ensemble methods significantly improve generation diversity, quality, and task-adaptive flexibility. Our work provides both theoretical foundations and practical guidelines for industrial-scale LLM selection, customization, and deployment.
Existing LLM ensemble methods often overlook model compatibility and rely on full-vocabulary probability alignment, resulting in low efficiency and high computational overhead. This work identifies model compatibility as a critical determinant of ensemble performance and proposes Union Top-k Ensemble—a novel paradigm that first selects a compatible subset of models and then performs lightweight probability aggregation solely over the union of each model’s top-k predicted tokens, thereby avoiding full-vocabulary alignment. Our approach establishes the first token-level, compatibility-aware ensemble framework that requires no retraining. Evaluated across multiple benchmark tasks, it significantly outperforms state-of-the-art ensemble methods in both accuracy and robustness while reducing computational cost by 30%–50%.
Time-series ensemble forecasting faces a critical trade-off between predictive accuracy and computational cost. This paper systematically evaluates ten base models and eight ensemble strategies on the M5 and VN1 retail datasets, measuring performance in point forecasting (RMSE) and probabilistic forecasting (CRPS), alongside computational overhead. Methodologically, we analyze ensemble size scalability, propose an “efficiency-driven ensemble” paradigm, and assess downsampling-based retraining frequency reduction. Key contributions: (1) Ensembles of only two to three models achieve near-optimal accuracy; (2) The efficiency-driven paradigm reduces average computational cost by over 40% while retaining ≥95% of baseline accuracy; (3) Reducing retraining frequency cuts training overhead by up to 70%, with negligible impact on point forecasts and robust performance in probabilistic forecasting. Results confirm that ensembling consistently improves prediction—especially probabilistic calibration—but high accuracy typically incurs high cost. Our framework delivers a scalable, cost-effective ensemble strategy for resource-constrained deployment.
This study investigates optimal parallel heterogeneous ensemble strategies for small- to medium-scale tabular classification tasks. Through systematic experiments on 56 OpenML CC18 datasets, the authors find that the instability of Blending and Stacking is mutually independent, and that Robust Soft Voting significantly outperforms Hard Voting in multi-class settings. Building on these insights, they propose a robust ensemble strategy and validate its effectiveness on the TabArena platform. The recommended approach significantly surpasses the Single Best method across 28 additional tasks and matches or exceeds the performance of all individual ensemble techniques considered.
This work addresses the lack of systematic understanding in efficiently fusing large language models (LLMs) fine-tuned with lightweight adapters in multi-task learning, particularly regarding the trade-offs among ensembling, merging, and routing strategies. The study systematically evaluates three parameter-efficient fusion approaches—output ensembling, parameter averaging, and input-dependent routing—and demonstrates that non-uniform fusion consistently outperforms uniform methods, with routing yielding significant performance gains despite its higher computational cost. To reconcile this efficiency–performance trade-off, the authors propose a low-overhead expert selection mechanism that combines clustering with greedy subset selection, achieving near-optimal performance while substantially reducing computational overhead, thereby striking an effective balance between model efficacy and efficiency.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
This work addresses the challenges of high computational costs and inefficient subset selection in large-scale image training. It proposes the SCOre-Stratified Selection (SCOSS) framework, which constructs a coreset through score-stratified sampling and integrates predictions from models trained on multiple independent subsets. By innovatively combining stratified sampling with ensemble learning, SCOSS significantly enhances the stability and generalization capability of coreset selection. Experimental results demonstrate that, across various sampling ratios, SCOSS-based coresets achieve state-of-the-art performance on the Simple Graph Convolution (SGC) model, surpassing support vector machines (SVMs) with only a small number of labeled samples while striking an excellent balance between accuracy and efficiency.
While ensemble models achieve strong performance, their large scale poses significant challenges in deployment, interpretability, and robustness verification. This work proposes PACE, a novel framework that uniquely integrates model compression and pruning through a two-stage alternating strategy. First, it enhances ensemble diversity by generating a set of diverse weak learners guided by a theoretically grounded active learning mechanism. Subsequently, it applies fidelity-aware pruning to the expanded ensemble, enabling controlled preservation of the original model’s behavioral characteristics. Extensive experiments demonstrate that PACE consistently outperforms existing compression and pruning methods across multiple tasks, achieving higher compression ratios while more faithfully retaining the behavior of the original ensemble.