🤖 AI Summary
This study investigates the value of geometric view synthesis under pretrained encoders for one-shot wine label retrieval. Methodologically, we construct a thousand-class one-shot retrieval benchmark and integrate LoRA fine-tuning with SAM-based localization to systematically compare geometric augmentation effects across DINO and SigLIP architectures. Experimental results demonstrate that fine-tuned DINO achieves up to a threefold accuracy improvement, whereas frozen SigLIP features yield only marginal gains. By quantifying the differential benefits of geometric synthesis across distinct architectures, this work delineates the effective applicability boundaries of the proposed approach, providing empirical evidence for view augmentation strategies in one-shot retrieval scenarios.
📝 Abstract
Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO vision transformer (ViT-S/16) recipe, geometric views raise top-1 accuracy from 34.1% to 62.6-63.7%, about three times the gain from two-dimensional (2D) augmentation. Frozen SigLIP 2-B already reaches 94.7%. A linear head over its frozen features gains 1.2-1.3 percentage points with the two geometric pipelines localized by the Segment Anything Model (SAM), while the other pipelines gain an inconclusive 0.3-0.6 points. Low-rank adaptation (LoRA) and validation-selected full fine-tuning show no clear gain within the reported confidence intervals; fixed-budget full fine-tuning loses 9-24 points. SAM localization supplies all six views for 99% of sources, compared with 43% for the edge-based front end. Recognition differences between the two cylinder constructions depend on the training recipe and are confounded by their crop and canvas conventions. Rendered-cylinder tests show different responses to source tilt, but an uncalibrated rim-ratio proxy establishes no corresponding trend in recognition on real photographs. An author-confirmed audit of 50 residual errors identifies 21 query-enrollment appearance mismatches, without establishing an irreducible error rate. These results support geometric synthesis for the tested self-supervised recipe and a smaller benefit through frozen-feature adaptation of the text-supervised encoder.