🤖 AI Summary
High-precision crop-weed semantic segmentation is critical for agricultural weeding robots, yet acquiring large-scale, pixel-accurate annotations from real-field environments remains prohibitively expensive. To address this, we propose a programmable synthetic image generation pipeline built on Blender, enabling efficient creation of diverse, photorealistic farmland scenes with precise pixel-level annotations by controlling key ecological and imaging parameters—including plant growth stages, weed density, illumination conditions, and camera viewpoints. Our method substantially mitigates domain shift: in sim-to-real transfer, it incurs only a 10% performance drop—outperforming state-of-the-art approaches. Remarkably, models trained solely on our synthetic data surpass those trained on real-world datasets in cross-domain segmentation tasks. This work presents the first systematic empirical validation of high-fidelity procedural synthetic data’s strong generalization capability in agricultural vision, establishing both theoretical grounding and practical feasibility for hybrid synthetic–real training paradigms.
📝 Abstract
Precise semantic segmentation of crops and weeds is necessary for agricultural weeding robots. However, training deep learning models requires large annotated datasets, which are costly to obtain in real fields. Synthetic data can reduce this burden, but the gap between simulated and real images remains a challenge. In this paper, we present a pipeline for procedural generation of synthetic crop-weed images using Blender, producing annotated datasets under diverse conditions of plant growth, weed density, lighting, and camera angle. We benchmark several state-of-the-art segmentation models on synthetic and real datasets and analyze their cross-domain generalization. Our results show that training on synthetic images leads to a sim-to-real gap of 10%, surpassing previous state-of-the-art methods. Moreover, synthetic data demonstrates good generalization properties, outperforming real datasets in cross-domain scenarios. These findings highlight the potential of synthetic agricultural datasets and support hybrid strategies for more efficient model training.