Score
Designs and analyzes pooling and aggregation mechanisms that convert variable-sized collections of instance-level features into a single bag- or sequence-level representation or prediction. This includes methods to weight, select, or suppress individual instances so salient but sparse or weak signals are emphasized and irrelevant/background instances are downweighted for robust classification or scoring.
Real-world data streams often originate from multiple underlying generative mechanisms, causing performance degradation in single-model approaches. To address this, we propose a mechanism-aware online ensemble learning framework. Its core is performance-driven data clustering: streaming instances are dynamically grouped based on similarity in predictive performance between features and targets; dedicated models are then trained in parallel for each mechanism-specific subgroup. Additionally, a dynamic weighting scheme updates model weights online using real-time validation errors. This framework departs from the conventional single-model assumption and introduces, for the first time, a predictive-performance-based clustering paradigm—enabling simultaneous mechanism identification, model specialization, and adaptive ensemble integration. Extensive experiments on multiple real-world streaming datasets demonstrate that our method significantly outperforms state-of-the-art single-model and static ensemble baselines, validating the effectiveness and generalization advantage of multi-mechanism modeling.
This work addresses the challenges of high computational costs and inefficient subset selection in large-scale image training. It proposes the SCOre-Stratified Selection (SCOSS) framework, which constructs a coreset through score-stratified sampling and integrates predictions from models trained on multiple independent subsets. By innovatively combining stratified sampling with ensemble learning, SCOSS significantly enhances the stability and generalization capability of coreset selection. Experimental results demonstrate that, across various sampling ratios, SCOSS-based coresets achieve state-of-the-art performance on the Simple Graph Convolution (SGC) model, surpassing support vector machines (SVMs) with only a small number of labeled samples while striking an excellent balance between accuracy and efficiency.
This paper addresses computational redundancy and diminishing returns arising from ensemble size expansion in data stream environments. It introduces, for the first time, a linear independence perspective on classifier voting to model ensemble performance. We establish linear independence as a fundamental mechanism for enhancing representational capacity and diversity, and derive a theoretical trade-off framework linking ensemble size to accuracy—yielding the minimal theoretical size required to achieve a target independence probability. Leveraging geometric modeling and weighted majority voting theory, we validate the framework empirically using OzaBagging and GOOWE. Experiments demonstrate that the framework accurately identifies performance saturation points for robust ensembles (e.g., OzaBagging), while revealing how high theoretical diversity may induce decision instability in less robust methods (e.g., GOOWE). The results provide a principled foundation for dynamic ensemble pruning and adaptive size control in streaming settings.
Standard Bagging ensembles often suffer from overconfidence and redundancy due to uniform voting weights that ignore the varying local competencies of base learners. This work proposes the SCSB framework, which unifies ensemble pruning and probability calibration into a joint optimization problem over the probability simplex. By minimizing out-of-bag loss with an added concave quadratic sparsity-inducing penalty, SCSB overcomes the theoretical limitation of the L1 norm—which fails to induce sparsity on the simplex—while preserving model-agnosticism. The method achieves compression rates up to 96%, substantially reduces expected calibration error, yields linear inference speedup, and maintains or even improves generalization accuracy in most cases.
This work addresses the computational complexity and high memory overhead associated with optimizing classifier weights in high- or infinite-dimensional kernel spaces. We propose a simple averaging classifier based on kernel mean embeddings and, for the first time, apply the herding algorithm to sparsify it. Unlike conventional weighted kernel classifiers, our approach constructs an unbiased, high-fidelity, and inherently parallelizable sparse approximation without altering the original learning objective. The resulting classifier preserves theoretical consistency and robustness while significantly reducing prediction latency and memory footprint. Empirically, it achieves competitive accuracy alongside superior efficiency and scalability. Moreover, its design naturally supports distributed implementation, offering a lightweight and reliable paradigm for large-scale kernel methods. (126 words)
This work addresses the problem of quantifying and enhancing the stability of machine learning models under feature perturbations to improve generalization. It introduces, for the first time, a measure termed Feature Instability (FI), which captures complementary generalization information relative to instance-level instability. Building upon algorithmic stability theory, the study analyzes a feature bagging mechanism and establishes theoretical guarantees—under both parametric linear and model-free settings—that this mechanism effectively reduces FI. Empirical results demonstrate that feature bagging significantly lowers FI, with only a few iterations needed to approach the stability achieved by infinite bagging; moreover, aggressive feature subsampling yields further improvements in stability.
This work addresses the challenge of effectively aggregating predictive distributions in deep ensembles to enhance both performance and reliability. From a log-likelihood perspective, the authors systematically analyze generalized mean-normalized aggregation and establish, for the first time, a unified theoretical framework that characterizes the behavior of aggregation across different orders \( r \). They rigorously prove that the aggregated prediction strictly outperforms any individual model if and only if \( r \in [0,1] \), thereby providing a solid theoretical foundation for the empirical success of linear pooling (\( r=1 \)) and geometric pooling (\( r \to 0 \)). Extensive experiments on image and text classification benchmarks confirm the practical relevance and effectiveness of the proposed theory.
Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.
This work addresses the challenge of prediction heterogeneity in high-dimensional multivariate time series forecasting, where global models often underperform and naive specialization risks negative transfer. The authors formulate adaptive pooling as a statistical decision problem and propose a validation-driven clustering framework that dynamically determines when and how to specialize sequences based on out-of-sample predictive performance rather than representational similarity. Clusters are iteratively refined using validation errors derived from Huber and pinball losses, while a leakage-free fallback strategy and a rigorous train-validation-test protocol ensure robustness. Evaluated on large-scale traffic datasets, the method significantly outperforms strong baselines and maintains stable performance even under weak heterogeneity.
This work addresses the degradation of model generalization in supervised learning caused by heterogeneity in training data. To mitigate this issue, the authors propose an input-space adaptive partitioning method grounded in the intrinsic heterogeneity of the data. By introducing a variance-based metric that quantifies the inconsistency in pairwise sample influence, they demonstrate that this variance is maximized under mixture distributions. Leveraging this property, the method automatically partitions the data into homogeneous subsets without requiring prior knowledge, enabling independent training of submodels on each subset. Experiments on EMNIST and synthetic datasets show significant improvements in test accuracy, confirming that the proposed variance metric effectively captures data heterogeneity and offers a novel pathway to enhance model generalization.