🤖 AI Summary
This study addresses the performance degradation of supervised learning on tabular data caused by latent label noise. To mitigate this issue, we propose a descriptor-driven adaptive ensemble framework. Specifically, the method constructs a data-centric inference module that integrates meta-models with confidence, neighborhood, and distribution detectors. By dynamically predicting the weights of these detectors, the framework enables automated data quality diagnosis and the selection of appropriate cleansing strategies. Experimental results demonstrate that the proposed approach achieves performance comparable to Confident Learning on standard benchmarks while significantly enhancing robustness across heterogeneous scenarios.
📝 Abstract
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.