🤖 AI Summary
This study addresses the scarcity of real-world data in domains such as construction safety monitoring by proposing a unified evaluation framework that systematically compares two synthetic data paradigms—Unity simulation and controllable diffusion models (CIA)—regarding their impact on object detection performance. The research reveals the performance bottlenecks inherent in single-source synthetic data and establishes a mixed-data balancing mechanism wherein moderate augmentation enhances generalization, whereas excessive synthetic data induces domain shift. Experimental results demonstrate that combining 90% real data with 10% Unity-simulated data improves mAP by 7.64% to 62.68%, while incorporating CIA-generated data further optimizes accuracy to 74.45%. These findings provide quantitative guidance for cost-effective data augmentation strategies.
📝 Abstract
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $Δ$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.