When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the value of geometric view synthesis under pretrained encoders for one-shot wine label retrieval. Methodologically, we construct a thousand-class one-shot retrieval benchmark and integrate LoRA fine-tuning with SAM-based localization to systematically compare geometric augmentation effects across DINO and SigLIP architectures. Experimental results demonstrate that fine-tuned DINO achieves up to a threefold accuracy improvement, whereas frozen SigLIP features yield only marginal gains. By quantifying the differential benefits of geometric synthesis across distinct architectures, this work delineates the effective applicability boundaries of the proposed approach, providing empirical evidence for view augmentation strategies in one-shot retrieval scenarios.
📝 Abstract
Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO vision transformer (ViT-S/16) recipe, geometric views raise top-1 accuracy from 34.1% to 62.6-63.7%, about three times the gain from two-dimensional (2D) augmentation. Frozen SigLIP 2-B already reaches 94.7%. A linear head over its frozen features gains 1.2-1.3 percentage points with the two geometric pipelines localized by the Segment Anything Model (SAM), while the other pipelines gain an inconclusive 0.3-0.6 points. Low-rank adaptation (LoRA) and validation-selected full fine-tuning show no clear gain within the reported confidence intervals; fixed-budget full fine-tuning loses 9-24 points. SAM localization supplies all six views for 99% of sources, compared with 43% for the edge-based front end. Recognition differences between the two cylinder constructions depend on the training recipe and are confounded by their crop and canvas conventions. Rendered-cylinder tests show different responses to source tilt, but an uncalibrated rim-ratio proxy establishes no corresponding trend in recognition on real photographs. An author-confirmed audit of 50 residual errors identifies 21 query-enrollment appearance mismatches, without establishing an irreducible error rate. These results support geometric synthesis for the tested self-supervised recipe and a smaller benefit through frozen-feature adaptation of the text-supervised encoder.
Problem

Research questions and friction points this paper is trying to address.

geometric view synthesis
wine label retrieval
one-shot learning
pretrained image encoders
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometric View Synthesis
One-Shot Retrieval
Vision-Language Models
Segment Anything Model
Self-Supervised Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yueh-Cheng Huang
Department of Computer Science and Information Engineering, National Dong Hwa University, Hualien 974301, Taiwan