🤖 AI Summary
Existing implicit ensemble methods struggle to control member diversity during training, limiting the performance and flexibility of uncertainty estimation. This work proposes σN-Ens, an implicit ensemble approach based on replicating normalization layers, where each ensemble member is modeled as a task within a multi-task architecture. Diversity is modulated through Sigmoid-bounded scalers applied to a shared backbone, and a Softmax temperature regularization term is introduced to govern the degree of parameter sharing among members. σN-Ens is the first method to enable controllable member diversity during training, introducing the notion of “modulated uncertainty” and leveraging temperature regularization to approach the accuracy–calibration Pareto frontier. Experiments show that σN-Ens matches or exceeds deep ensembles on CIFAR-10/100, ImageNet, and SST-2 with substantially lower parameter overhead, exhibits robust calibration under distribution shift, and consistently improves with larger ensemble sizes.
📝 Abstract
Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit ensembles lower this cost by sharing a single backbone across members. Member diversity is a primary determinant of ensemble quality, yet no implicit ensemble can shape it during training; existing methods fix it at initialisation or build it into the architecture. We introduce $σ$N-Ens, a normalisation-based implicit ensemble that treats each member as a task in a multi-task architecture and modulates the shared backbone through sigmoid-bounded scalers. We also introduce a softmax-temperature regulariser, which shapes the equilibrium level of sharing between members and traces the accuracy-calibration frontier. Because only normalisation layers are replicated, the mechanism can wrap convolutional and transformer backbones alike, also allowing pretrained models to be adapted through a short fine-tune. We frame the epistemic uncertainty such an ensemble expresses as modulation uncertainty, and explain why its calibration holds under input corruption, and why its out-of-distribution detection is weaker. Our method is evaluated across ResNets and transformers on CIFAR-10/100, ImageNet and SST-2. $σ$N-Ens matches or outperforms deep ensembles at a fraction of their parameter cost, scales with ensemble size where partitioning methods collapse, and maintains calibration under distribution shift.