Score
Designs and implements methods that combine multiple models' probabilistic outputs (class posteriors) into a single robust multiclass probability estimate using soft/majority-vote fusion—e.g., weighted averaging, confidence-aware aggregation, or probabilistic majority voting—applied to heterogeneous or routed base learners. Builds and evaluates these ensembles for calibration, stability across instances or templates, resistance to noisy or adversarial predictions, and overall accuracy improvements over single-model or hard-vote baselines, including techniques for outlier rejection and reliability weighting.
This study investigates optimal parallel heterogeneous ensemble strategies for small- to medium-scale tabular classification tasks. Through systematic experiments on 56 OpenML CC18 datasets, the authors find that the instability of Blending and Stacking is mutually independent, and that Robust Soft Voting significantly outperforms Hard Voting in multi-class settings. Building on these insights, they propose a robust ensemble strategy and validate its effectiveness on the TabArena platform. The recommended approach significantly surpasses the Single Best method across 28 additional tasks and matches or exceeds the performance of all individual ensemble techniques considered.
To address the unreliability of prediction confidence in deep neural network image classification, this paper proposes a lightweight meta-model classifier ensemble method that achieves efficient uncertainty calibration without requiring additional calibration data. The approach constructs a parameter-efficient meta-model to fuse outputs from multiple base classifiers and jointly evaluates calibration performance using majority voting alongside Expected Calibration Error (ECE) and Maximum Calibration Error (MCE). For the first time, both theoretical analysis and empirical evaluation demonstrate its significant calibration advantages: across diverse mainstream CNN architectures, it reduces ECE and MCE by over 40% on average while preserving classification accuracy nearly unchanged; moreover, its parameter count is 3–5× smaller than conventional model ensembles. This work establishes a novel paradigm for high-reliability, low-overhead model calibration.
This work addresses the limited generalization improvement and uncontrollable error cancellation in Bayesian neural network ensembles for classification, stemming from neglecting inter-model error correlations. We propose a weighted ensemble method optimized via a second-order PAC-Bayes bound. Its core innovation is the first explicit incorporation of error correlation into a PAC-Bayesian weighting framework, coupled with a tandem loss for robust fusion. We provide a rigorous, non-vacuous generalization error bound grounded in PAC-Bayes theory. Empirically, our method significantly outperforms standard Bayesian ensembling in both accuracy and generalization calibration, while matching or exceeding uniformly weighted ensembles. Moreover, it enables safe fusion of multiple checkpoint models from a single training run—ensuring both theoretical guarantees and practical applicability.
This work investigates the interplay between certified robustness and generalization of smoothed majority-voting classifiers within the PAC-Bayesian framework. Existing approaches suffer from dimension-dependent guarantees and struggle to jointly optimize robustness and accuracy. To address this, we propose the first dimension-agnostic spectral regularization strategy for smoothed voting by incorporating the spectral norm of classifier weights into the architecture. Coupled with spherical Gaussian input smoothing during training, we derive tight, unified theoretical upper bounds that simultaneously characterize both the certified robust radius and the generalization error. Experiments on standard benchmarks—including CIFAR-10—demonstrate that our method significantly improves certified accuracy (+2.3%) while preserving natural accuracy (±0.5%), thereby achieving an effective trade-off between robustness and generalization.
This paper addresses computational redundancy and diminishing returns arising from ensemble size expansion in data stream environments. It introduces, for the first time, a linear independence perspective on classifier voting to model ensemble performance. We establish linear independence as a fundamental mechanism for enhancing representational capacity and diversity, and derive a theoretical trade-off framework linking ensemble size to accuracy—yielding the minimal theoretical size required to achieve a target independence probability. Leveraging geometric modeling and weighted majority voting theory, we validate the framework empirically using OzaBagging and GOOWE. Experiments demonstrate that the framework accurately identifies performance saturation points for robust ensembles (e.g., OzaBagging), while revealing how high theoretical diversity may induce decision instability in less robust methods (e.g., GOOWE). The results provide a principled foundation for dynamic ensemble pruning and adaptive size control in streaming settings.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
To address the insufficient robustness of federated inference—particularly its vulnerability to malicious attacks causing prediction inaccuracies—this paper formally defines the robust federated inference task and proposes a nonlinear aggregation modeling framework grounded in adversarial machine learning. Methodologically, it introduces a DeepSet-based joint training-inference defense mechanism that integrates adversarial training with test-time robust aggregation, enabling localized model deployment and privacy preservation. A key contribution is the unified security analysis of both linear and nonlinear aggregators, yielding provably robust aggregation optimization. Extensive evaluation on multiple benchmark datasets demonstrates that the proposed method improves accuracy by 4.7–22.2 percentage points over state-of-the-art robust aggregation schemes, significantly enhancing system resilience against adversarial attacks and overall reliability.
To address the insufficient robustness of weapon detection under challenging conditions—including occlusion, illumination variations, and cluttered backgrounds—this paper proposes a detection framework based on ensemble SSD models with diverse backbone networks. The method jointly emphasizes model diversity and discovery-aware fusion: it constructs heterogeneous SSD variants using VGG16, ResNet50, EfficientNet, and MobileNetV3 as backbones, and integrates their predictions via a weighted box fusion (WBF) mechanism that employs “max”-confidence weighting. Evaluated on a multi-class weapon dataset, the ensemble achieves an mAP of 0.838—outperforming the best single-model baseline by 2.95% and surpassing conventional fusion approaches. Moreover, it demonstrates consistent performance gains across all challenging scenarios, confirming its effectiveness and generalizability in real-world weapon detection tasks.
This work addresses the significant performance degradation of deep learning models on tail classes in long-tailed imbalanced datasets by introducing, for the first time, a Bayesian perspective to reveal and quantify the model’s “preference bias” toward head classes. Leveraging Bayesian theory and the assumption that class frequencies follow a power-law distribution, the authors propose a computationally efficient and easily integrable log-adjustment method to achieve class balance during training. A novel metric is introduced to evaluate preference bias, and extensive experiments on multiple large-scale long-tailed benchmarks demonstrate that the proposed approach substantially outperforms existing techniques—such as Balanced Softmax—while exhibiting strong scalability and practical utility.