Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

📅 2026-08-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of scarce training data and poor cross-scenario generalization in small object detection under heavy vegetation occlusion. It proposes a novel annotation-free synthetic data generation method that leverages a small set of unlabeled in-the-wild images to reconstruct 3D vegetated scenes guided by a vision-language model (VLM). Realistic annotated images are synthesized by embedding 3D object meshes into these scenes, followed by controllable diffusion-based inpainting to refine object textures and occlusion patterns. A hierarchical mask-locking mechanism is introduced to enhance recall for minority classes while preserving semantic consistency. Evaluated on a humanitarian demining benchmark, detectors trained solely on this synthetic data match or even surpass models trained with substantially more real annotated data from different domains, demonstrating significantly improved unsupervised cross-domain adaptability.
📝 Abstract
Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing labels collected at another site, we synthesize labeled training images from a handful of unlabeled photographs of the deployment site itself. A vision--language model generates a coarse 3D vegetation scene from one photograph; placing 3D object meshes in the scene yields bounding boxes, segmentation masks, and per-instance occlusion directly from the scene geometry, without manual annotation. A lightweight adapter fine-tuned on the photographs conditions a diffusion pass that re-textures the renders, and a graded mask-lock sets how much diffusion may touch the object itself. In our runs this grade was the most influential curation choice: lightly diffusing the object improves minority-class recall over fully protecting its pixels, while unrestricted diffusion dissolves it. Trained on these images, a standard detector matched or exceeded its counterpart trained on a larger labeled dataset of real images from a different site, consistently across seeds on a humanitarian-demining benchmark; the comparison is thus unsupervised site adaptation from a handful of photographs against conventional cross-site label reuse. In our ablations the gains were largely insensitive to the photograph and crop budgets, and in-domain accuracy did not predict cross-site performance.
Problem

Research questions and friction points this paper is trying to address.

object detection
vegetation occlusion
cross-site generalization
labeled data scarcity
domain adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
3D scene synthesis
diffusion re-texturing
mask-lock mechanism
unsupervised domain adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mario Malizia
Royal Military Academy, Brussels, Belgium
M
Marnix Enting
Royal Military Academy, Brussels, Belgium
R
Rob Haelterman
Royal Military Academy, Brussels, Belgium
K
Ken Hasselmann
Royal Military Academy, Brussels, Belgium