🤖 AI Summary
This study addresses the challenges of counting and localization in agricultural scenarios, where target appearance, scale, and density vary significantly. We propose a parameter-efficient exemplar-guided framework that leverages DINOv3 multi-scale feature conditioning and progressive point decoding to jointly perform counting and localization using frozen features, eliminating the need for category-specific retraining. A novel missed-detection recovery mechanism is introduced to augment supervision signals, alongside an exemplar-adaptive non-maximum suppression strategy to filter redundant predictions, substantially reducing model parameters. Experimental results demonstrate that the proposed method achieves a mean absolute error (MAE) of 11.92 on the TPC-268 benchmark, representing a 9.7% error reduction, and attains a zero-shot MAE of 14.25 on FSC-147, outperforming existing state-of-the-art approaches.
📝 Abstract
Accurate counting and localization of plants and their organs support phenotyping and yield estimation, yet target appearance, scale, and density vary widely across species and imaging conditions. Exemplar boxes specify the target without category-specific retraining, and point predictions identify the individual instances contributing to the count. We introduce AgriCountDINO, a parameter-efficient exemplar-guided framework for joint counting and localization. It conditions frozen multiscale DINOv3 features on exemplar appearance and size, then progressively decodes them into target points. Missed-object recovery extends supervision to targets overlooked by initial matching, and exemplar-adaptive point NMS filters duplicate predictions according to exemplar scale. With 8.4M trainable parameters, approximately one-tenth of TasselNetV4's, AgriCountDINO achieves a three-shot MAE of 11.92 on the TPC-268 benchmark, reducing counting error by 9.7\% while providing individual target locations. Trained only on TPC-268, it achieves a zero-shot MAE of 14.25 on unseen generic object categories in FSC-147, improving upon the best compared zero-shot method by 6.0\% without target-domain training or fine-tuning.