Hard, Yet Reducible: Controlled Forward Transfer for Synthetic Degradation Curation

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of evaluating the training utility of synthetically degraded data for dense prediction tasks, where assessments are often confounded by performance degradation on clean images. To overcome this, we propose a controlled and Reversible Degradation Gap (RDG) metric alongside a corresponding data filtering method. The core innovation lies in introducing a budget-matched dual-probe mechanism that effectively isolates clean-domain performance bias. Furthermore, we define a correctable severity band that dynamically adapts to model architectures and computational budgets, enabling the precise selection of high-value synthetic samples. Experimental results demonstrate that our approach significantly improves semantic segmentation and salient object detection performance under equivalent computational budgets. Notably, it remains compatible with existing generative pipelines while preserving accuracy on clean images.
πŸ“ Abstract
Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbf{controlled Reducible Degradation Gap} (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbf{Curation of Reducible Bands} (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method{} improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

synthetic degradation
dense prediction
data curation
training utility
generalization gain
Innovation

Methods, ideas, or system contributions that make the work stand out.

Controlled Reducible Degradation Gap
Synthetic Degradation Curation
Dense Prediction
Data Selection
Training Budget