Fitting Multiple Machine Learning Models with Performance Based Clustering

๐Ÿ“… 2024-11-10
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Real-world data streams often originate from multiple underlying generative mechanisms, causing performance degradation in single-model approaches. To address this, we propose a mechanism-aware online ensemble learning framework. Its core is performance-driven data clustering: streaming instances are dynamically grouped based on similarity in predictive performance between features and targets; dedicated models are then trained in parallel for each mechanism-specific subgroup. Additionally, a dynamic weighting scheme updates model weights online using real-time validation errors. This framework departs from the conventional single-model assumption and introduces, for the first time, a predictive-performance-based clustering paradigmโ€”enabling simultaneous mechanism identification, model specialization, and adaptive ensemble integration. Extensive experiments on multiple real-world streaming datasets demonstrate that our method significantly outperforms state-of-the-art single-model and static ensemble baselines, validating the effectiveness and generalization advantage of multi-mechanism modeling.

Technology Category

Machine Learning: Ensemble MethodsMultiagent Systems: Mechanism DesignData Mining & Knowledge Management: Data Stream Mining

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
๐Ÿ“ Abstract
Traditional machine learning approaches assume that data comes from a single generating mechanism, which may not hold for most real life data. In these cases, the single mechanism assumption can result in suboptimal performance. We introduce a clustering framework that eliminates this assumption by grouping the data according to the relations between the features and the target values and we obtain multiple separate models to learn different parts of the data. We further extend our framework to applications having streaming data where we produce outcomes using an ensemble of models. For this, the ensemble weights are updated based on the incoming data batches. We demonstrate the performance of our approach over the widely-studied real life datasets, showing significant improvements over the traditional single-model approaches.
Problem

Research questions and friction points this paper is trying to address.

Non-uniform data types
Model selection
Continuous data stream
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Learning
Non-uniform Data Handling
Dynamic Model Adaptation
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Bilkent University
M
Mehmet E. Lorasdagi
Department of Electrical and Electronics Engineering, Bilkent University, Turkey
A
A. B. Koc
Department of Electrical and Electronics Engineering, Bilkent University, Turkey
A
Ali T. Koc
Department of Electrical and Electronics Engineering, Bilkent University, Turkey
S
Suleyman S. Kozat
Department of Electrical and Electronics Engineering, Bilkent University, Turkey