Score
Designs, trains, and evaluates ensembles of deep neural networks that aggregate predictions from multiple models, including methods for combining member outputs and matching an ensemble’s parameter budget to baseline models. This competence covers selecting architectures and training procedures to promote member diversity and calibration, and analyzing ensemble robustness and performance under missing, imbalanced, or varying member inputs.
Existing ensemble methods (e.g., greedy or random ensembling) employ static, sample-agnostic base-model weights, limiting expressive capacity and generalization. This paper proposes the Dynamic Posterior Neural Ensemble (DPNE), the first framework to introduce Random Prediction Dropout (RPD)—a theoretically grounded diversity regularization mechanism that provably enhances base-model diversity under a derived lower bound. DPNE employs a neural architecture to learn sample-wise adaptive weights end-to-end, eliminating restrictive pre-specified weight structures. Extensive experiments across CV, NLP, and tabular tasks demonstrate that DPNE significantly outperforms state-of-the-art ensemble baselines, effectively mitigating overfitting while improving both in-distribution accuracy and out-of-distribution robustness. The core contributions are: (1) dynamic posterior weight modeling conditioned on input samples, and (2) theoretically guaranteed diversity regularization via RPD.
Single-architectural vision models face inherent performance bottlenecks in image classification. Method: This paper proposes a cross-paradigm ensemble framework that preserves architectural integrity by integrating three heterogeneous architectures—CNN (ResNet), MLP-Mixer, and Vision Transformer—via weighted averaging and majority voting, without architectural modification or feature-space alignment. Crucially, the framework implicitly isolates their respective feature spaces while enabling synergistic gains. Contribution/Results: It introduces the first complementary analysis paradigm grounded in architectural orthogonality. Evaluated end-to-end on ImageNet, the ensemble surpasses prior single-model SOTA in top-1 accuracy while reducing overall inference latency—establishing a new benchmark for efficient, high-accuracy image classification.
To address the unreliability of prediction confidence in deep neural network image classification, this paper proposes a lightweight meta-model classifier ensemble method that achieves efficient uncertainty calibration without requiring additional calibration data. The approach constructs a parameter-efficient meta-model to fuse outputs from multiple base classifiers and jointly evaluates calibration performance using majority voting alongside Expected Calibration Error (ECE) and Maximum Calibration Error (MCE). For the first time, both theoretical analysis and empirical evaluation demonstrate its significant calibration advantages: across diverse mainstream CNN architectures, it reduces ECE and MCE by over 40% on average while preserving classification accuracy nearly unchanged; moreover, its parameter count is 3–5× smaller than conventional model ensembles. This work establishes a novel paradigm for high-reliability, low-overhead model calibration.
Under hardware resource constraints, compact deep ensembles suffer significant degradation in accuracy, calibration, uncertainty estimation, and out-of-distribution (OOD) detection. To address this, we propose Packed-Ensembles (PE): a memory-aligned, lightweight structured ensemble method that enables backbone sharing and single-pass forward propagation within the memory budget of a single model, achieved via spatial-dimension modulation encoding and grouped convolutions. PE is the first approach to compress deep ensembles into a compact, end-to-end trainable structured packing—preserving ensemble diversity and statistical performance while drastically reducing computational and memory overhead. Experiments demonstrate that PE matches standard deep ensembles in accuracy, calibration, OOD detection, and robustness to distributional shift, while accelerating both training and inference. Crucially, its memory footprint remains strictly bounded by that of a single base model.
This study addresses the critical challenge of efficiently obtaining reliable model uncertainty under resource-constrained and low-latency conditions. It presents a systematic evaluation of BatchEnsemble in terms of accuracy, calibration, and out-of-distribution (OOD) detection performance. Through comprehensive empirical analyses—including comparisons with deep ensembles, calibration assessments, and controlled investigations of functional and parameter-space similarity among ensemble members on MNIST—the work reveals, for the first time, that BatchEnsemble members exhibit high homogeneity and lack diversity. Results show that BatchEnsemble performs comparably to a single model on CIFAR-10, CIFAR-10-C, and SVHN, while its members on MNIST are nearly identical, failing to capture the predictive diversity characteristic of true ensembles. These findings cast doubt on BatchEnsemble’s effectiveness as an efficient ensemble method.
This paper addresses the joint optimization of weight decay, temperature scaling, and early stopping in deep ensemble learning to simultaneously improve predictive accuracy and uncertainty calibration. To overcome evaluation bias and suboptimal data utilization inherent in conventional independent hyperparameter tuning, we propose Partial Overlap Validation—a novel cross-validation strategy that enables valid joint assessment while maximizing training data usage. Our method integrates joint regularization, learnable temperature scaling, and adaptive early stopping within an enhanced cross-validation framework. Experiments across multiple benchmark tasks demonstrate that joint optimization consistently outperforms independent tuning: Expected Calibration Error (ECE) decreases by up to 32%, and Negative Log-Likelihood (NLL) improves significantly. Moreover, our validation strategy explicitly reveals the fundamental trade-off between individual and joint hyperparameter optimization. The implementation is publicly available.
Neural ensemble search faces significant challenges due to the exponential growth of the joint search space over individual architectures and ensemble compositions, rendering brute-force approaches computationally infeasible. To address this, this work proposes a bi-objective surrogate modeling approach that introduces, for the first time, a co-optimization mechanism: two independently trained surrogate models estimate both the predictive accuracy and diversity potential of candidate architectures. Leveraging a directed acyclic graph representation of architectures, the method efficiently guides the ensemble search process. Experimental results on FashionMNIST, CIFAR-10, and CIFAR-100 demonstrate that the resulting ensembles achieve performance comparable to or better than established baselines such as Deep Ensembles and random search, while substantially reducing computational complexity.
This work addresses the open question of whether, under a fixed parameter budget, model capacity should be concentrated in a single wide path or distributed across multiple narrow branches. The authors propose the Multi-Narrow transformation, which restructures standard CNNs into ensembles of independent narrow pathways, and conduct a systematic evaluation comparing single-wide and multi-narrow architectures under comparable total parameter counts across varying data regimes, network designs, and datasets. Experimental results demonstrate that Multi-Narrow significantly outperforms baseline single-wide models in low-data settings, owing to reduced inter-path feature redundancy and enhanced diversity, which collectively improve generalization. However, it underperforms when ample training data is available. The study delineates the effective operating regime of multi-narrow structures and elucidates their underlying mechanisms.
This work addresses the trade-off between computational cost and accuracy in neural network ensembles, where existing ensemble methods are computationally expensive while conventional weight aggregation techniques often sacrifice performance. To bridge this gap, the authors propose a “partial fusion” framework that formulates weight aggregation as a generalized pruning process. By leveraging partial optimal transport, the method matches and fuses the most similar neurons across models, allowing for neuron deletion, isolation, or linear combination. A similarity metric at the neuron level enables a controllable balance between computational overhead and model accuracy. Experiments demonstrate that partial fusion significantly reduces computational costs while preserving accuracy close to that of full ensembles; furthermore, its single-model variant outperforms traditional pruning approaches.
Existing implicit ensemble methods struggle to control member diversity during training, limiting the performance and flexibility of uncertainty estimation. This work proposes σN-Ens, an implicit ensemble approach based on replicating normalization layers, where each ensemble member is modeled as a task within a multi-task architecture. Diversity is modulated through Sigmoid-bounded scalers applied to a shared backbone, and a Softmax temperature regularization term is introduced to govern the degree of parameter sharing among members. σN-Ens is the first method to enable controllable member diversity during training, introducing the notion of “modulated uncertainty” and leveraging temperature regularization to approach the accuracy–calibration Pareto frontier. Experiments show that σN-Ens matches or exceeds deep ensembles on CIFAR-10/100, ImageNet, and SST-2 with substantially lower parameter overhead, exhibits robust calibration under distribution shift, and consistently improves with larger ensemble sizes.