Effects of Training Data Quality on Classifier Performance

📅 2026-02-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The impact of training data quality on classifier performance is often overlooked. In the context of metagenomic DNA sequence assembly, this study systematically evaluates the behavior of Bayesian classifiers, neural networks, partition models, and random forests under various training data degradation scenarios. The findings reveal that as data quality deteriorates, all classifiers exhibit a “catastrophic” degradation pattern—shifting from substantially correct predictions to essentially random guesses. Concurrently, decision boundaries become sparser, and inter-classifier agreement paradoxically increases, indicating a convergence in error patterns under low-quality training conditions. This work provides the first quantitative characterization of the relationship between data quality and heterogeneity in classifier behavior, offering new insights for designing robust classification systems in data-scarce or noisy environments.

Technology Category

Machine Learning: Multi-class/Multi-label Learning & Extreme ClassificationSearch and Optimization: Metareasoning and MetaheuristicsData Mining & Knowledge Management: Anomaly/Outlier Detection

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Web data quality in the era of algorithmically-generated contentGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
We describe extensive numerical experiments assessing and quantifying how classifier performance depends on the quality of the training data, a frequently neglected component of the analysis of classifiers. More specifically, in the scientific context of metagenomic assembly of short DNA reads into "contigs," we examine the effects of degrading the quality of the training data by multiple mechanisms, and for four classifiers -- Bayes classifiers, neural nets, partition models and random forests. We investigate both individual behavior and congruence among the classifiers. We find breakdown-like behavior that holds for all four classifiers, as degradation increases and they move from being mostly correct to only coincidentally correct, because they are wrong in the same way. In the process, a picture of spatial heterogeneity emerges: as the training data move farther from analysis data, classifier decisions degenerate, the boundary becomes less dense, and congruence increases.
Problem

Research questions and friction points this paper is trying to address.

training data quality
classifier performance
metagenomic assembly
data degradation
classifier congruence
Innovation

Methods, ideas, or system contributions that make the work stand out.

training data quality
classifier performance
metagenomic assembly
spatial heterogeneity
classifier congruence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alan F. Karr
Department of Statistics, Operations and Data Science, Temple University, Philadelphia PA 19122 and Fraunhofer USA Center Mid-Atlantic, Riverdale MD 20737
R
Regina Ruane
Computational Social Science Laboratory, University of Pennsylvania, Philadelphia, PA 19104