K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the degradation of original feature importance rankings caused by conventional oversampling methods on imbalanced tabular data. To preserve interpretability, the authors propose a generate-then-filter oversampling framework that explicitly constrains the Kendall’s tau correlation between the feature importance rankings of synthetic and original samples, with stronger protection afforded to highly important features. The framework adopts a generator-agnostic “generate-and-filter” strategy, making it compatible with diverse model interpretation techniques. Experimental results across 20 datasets demonstrate that the proposed method significantly improves fidelity in feature importance preservation, enhances explanation consistency, and boosts class separability, all while achieving better predictive performance without incurring prohibitive computational overhead.
📝 Abstract
Oversampling is widely used to address class imbalance in tabular classification, but existing methods can distort the feature importance ranking underlying model explanations. Although recent studies have quantified this distortion by comparing real and synthetic data, none have actively sought to prevent it. In this paper, we introduce Kendall-constrained Importance-Preserving Oversampling (K-IPO), a generator-agnostic, "generate-then-select" framework that preserves the original data's feature importance ranking during augmentation. K-IPO iteratively generates minority-class candidates and accepts them only if their inclusion maintains a user-defined minimum Kendall's tau (τ) correlation with the reference ranking. Optionally, stricter constraints can be applied to the highest-ranked features. We evaluated K-IPO on 20 imbalanced binary classification datasets using three classifiers and multiple explanation methods. In most cases, K-IPO achieved the best or tied-best results in feature importance preservation, explanation consistency, and class separability. It also generally improved predictive performance while maintaining competitive computational overhead.
Problem

Research questions and friction points this paper is trying to address.

class imbalance
oversampling
feature importance
Kendall's tau
tabular data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kendall's tau
feature importance preservation
oversampling
imbalanced tabular data
explanation consistency
🔎 Similar Papers
M
Marios Tyrovolas
Department of Informatics and Telecommunications, University of Ioannina, Arta, Greece
A
Argiris Sofotasios
Department of Computer Engineering and Informatics, University of Patras, Patras, Greece
D
Dimitris Metaxakis
Department of Computer Engineering and Informatics, University of Patras, Patras, Greece
G
Georgios Mermigkis
Industrial Systems Institute, Athena Research Center, Patras, Greece
George Georgoulas
George Georgoulas
Post Doc researcher
Machine LearningData MiningFault Detection-DiagnosisDecision Support SystemsSignal Processing
Panagiotis Hadjidoukas
Panagiotis Hadjidoukas
University of Patras
Parallel and Distributed Computing
Chrysostomos Stylios
Chrysostomos Stylios
Industrial Systems Institute/ Athena RC & University of Ioannina
Fuzzy Cognitive MapsBig DataComputational IntelligenceBiosignal processing & analysisModeling and Decision Support Syste