๐ค AI Summary
Real-world data streams often originate from multiple underlying generative mechanisms, causing performance degradation in single-model approaches. To address this, we propose a mechanism-aware online ensemble learning framework. Its core is performance-driven data clustering: streaming instances are dynamically grouped based on similarity in predictive performance between features and targets; dedicated models are then trained in parallel for each mechanism-specific subgroup. Additionally, a dynamic weighting scheme updates model weights online using real-time validation errors. This framework departs from the conventional single-model assumption and introduces, for the first time, a predictive-performance-based clustering paradigmโenabling simultaneous mechanism identification, model specialization, and adaptive ensemble integration. Extensive experiments on multiple real-world streaming datasets demonstrate that our method significantly outperforms state-of-the-art single-model and static ensemble baselines, validating the effectiveness and generalization advantage of multi-mechanism modeling.
๐ Abstract
Traditional machine learning approaches assume that data comes from a single generating mechanism, which may not hold for most real life data. In these cases, the single mechanism assumption can result in suboptimal performance. We introduce a clustering framework that eliminates this assumption by grouping the data according to the relations between the features and the target values and we obtain multiple separate models to learn different parts of the data. We further extend our framework to applications having streaming data where we produce outcomes using an ensemble of models. For this, the ensemble weights are updated based on the incoming data batches. We demonstrate the performance of our approach over the widely-studied real life datasets, showing significant improvements over the traditional single-model approaches.