π€ AI Summary
This study addresses the challenge of evaluating the training utility of synthetically degraded data for dense prediction tasks, where assessments are often confounded by performance degradation on clean images. To overcome this, we propose a controlled and Reversible Degradation Gap (RDG) metric alongside a corresponding data filtering method. The core innovation lies in introducing a budget-matched dual-probe mechanism that effectively isolates clean-domain performance bias. Furthermore, we define a correctable severity band that dynamically adapts to model architectures and computational budgets, enabling the precise selection of high-value synthetic samples. Experimental results demonstrate that our approach significantly improves semantic segmentation and salient object detection performance under equivalent computational budgets. Notably, it remains compatible with existing generative pipelines while preserving accuracy on clean images.
π Abstract
Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbf{controlled Reducible Degradation Gap} (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbf{Curation of Reducible Bands} (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method{} improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.