A Data-Centric Perspective on the Influence of Image Data Quality in Machine Learning Models

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Image data quality significantly impacts model performance, yet systematic evaluation methodologies remain lacking. This paper proposes an automated image quality assessment pipeline integrating CleanVision and Fastdup, specifically targeting quantitative detection of degradation artifacts—including blur and excessive scaling. We introduce a novel automatic threshold selection mechanism that robustly identifies images with compromised critical visual features without manual parameter tuning, and enhance near-duplicate sample deduplication. Under a binary classification evaluation framework, our method achieves an F1-score of 0.9468 (+0.2674) for single-distortion detection and 0.8557 (+0.1110) for dual-distortion detection; near-duplicate detection F1 improves to 0.7928 (+0.3352). These results demonstrate substantial gains in both accuracy and generalizability of data cleaning.

Technology Category

Computer Vision: Adversarial Attacks & RobustnessMachine Learning: Multi-instance/Multi-view LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Web Mining and Content Analysis: Web data integration and cleaningSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
In machine learning, research has traditionally focused on model development, with relatively less attention paid to training data. As model architectures have matured and marginal gains from further refinements diminish, data quality has emerged as a critical factor. However, systematic studies on evaluating and ensuring dataset quality in the image domain remain limited. This study investigates methods for systematically assessing image dataset quality and examines how various image quality factors influence model performance. Using the publicly available and relatively clean CIFAKE dataset, we identify common quality issues and quantify their impact on training. Building on these findings, we develop a pipeline that integrates two community-developed tools, CleanVision and Fastdup. We analyze their underlying mechanisms and introduce several enhancements, including automatic threshold selection to detect problematic images without manual tuning. Experimental results demonstrate that not all quality issues exert the same level of impact. While convolutional neural networks show resilience to certain distortions, they are particularly vulnerable to degradations that obscure critical visual features, such as blurring and severe downscaling. To assess the performance of existing tools and the effectiveness of our proposed enhancements, we formulate the detection of low-quality images as a binary classification task and use the F1 score as the evaluation metric. Our automatic thresholding method improves the F1 score from 0.6794 to 0.9468 under single perturbations and from 0.7447 to 0.8557 under dual perturbations. For near-duplicate detection, our deduplication strategy increases the F1 score from 0.4576 to 0.7928. These results underscore the effectiveness of our workflow and provide a foundation for advancing data quality assessment in image-based machine learning.
Problem

Research questions and friction points this paper is trying to address.

Systematically assessing image dataset quality for machine learning models
Investigating how various image quality factors influence model performance
Developing automated tools to detect problematic images without manual tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematically assesses image dataset quality using CleanVision and Fastdup tools
Introduces automatic threshold selection for detecting problematic images
Enhances near-duplicate detection with improved deduplication strategy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pei-Han Chen
Department of Applied Mathematics, National Sun Yat-sen University, Kaohsiung, Taiwan
S
Szu-Chi Chung
Department of Applied Mathematics, National Sun Yat-sen University, Kaohsiung, Taiwan