From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of real-world data in domains such as construction safety monitoring by proposing a unified evaluation framework that systematically compares two synthetic data paradigms—Unity simulation and controllable diffusion models (CIA)—regarding their impact on object detection performance. The research reveals the performance bottlenecks inherent in single-source synthetic data and establishes a mixed-data balancing mechanism wherein moderate augmentation enhances generalization, whereas excessive synthetic data induces domain shift. Experimental results demonstrate that combining 90% real data with 10% Unity-simulated data improves mAP by 7.64% to 62.68%, while incorporating CIA-generated data further optimizes accuracy to 74.45%. These findings provide quantitative guidance for cost-effective data augmentation strategies.
📝 Abstract
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $Δ$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.
Problem

Research questions and friction points this paper is trying to address.

Object Detection
Data Scarcity
Synthetic Data Augmentation
Dataset Balance
Domain Drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion-based augmentation
Unity simulation
Dataset balance
Object detection
Domain drift
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Mohamed Benkedadra
Mohamed Benkedadra
Université de Mons
Machine LearningDeep LearningComputer Vision
A
Aissa Saoudi
Université Polytechnique Hauts-de-France
M
Maxime Gloesener
DeepILIA, ILIA, Université de Mons
S
Sidi Ahmed Mahmoudi
DeepILIA, ILIA, Université de Mons
Matei Mancas
Matei Mancas
Numediart Institute for Creative Technologies, University of Mons (UMONS)
visual attentionsaliencycreative technologies