🤖 AI Summary
This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.
📝 Abstract
Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation is expensive, imprecise, or infeasible. In this article we argue that the central question for specialized vision systems is not how to generate more data, but which data to generate. We therefore discuss curated synthetic data, whose scene content, appearance variations, sensing characteristics, and annotations are deliberately designed around a given task. We examine three complementary curation paradigms. Procedural rendering offers explicit control over scene parameters and the annotations follow by construction. Physically-based simulation encodes the mechanism behind an observed effect and yields exactly aligned supervision pairs. Generative AI learns sensor-specific appearance from small real seed sets and attains plausible realism, though it remains prone to hallucination and to inaccurate annotation. These paradigms are illustrated with examples from industrial surface defect detection, restoration of degraded digitized autochrome plates, and 6DoF pose estimation from RGB and event data. Using these examples, we analyze different data regimes and training strategies that combine synthetic and real data across these paradigms. We conclude that curated synthetic data are best understood as a complement to real observations, and that hybrid pipelines combining controllable supervision with learned appearance are the most promising direction for reliable sim-to-real transfer.