🤖 AI Summary
This study addresses the challenge that small-object features in vision-language encoders are prone to confusion with background, hindering accurate localization. To this end, it proposes a prototype-contrastive and local-magnification mechanism built upon a frozen CLIP backbone. Foreground–background prototypes are constructed from the support set, and precise localization is achieved through densely sampled query windows, reprojection, and coverage averaging. The per-class memory footprint remains fixed, enabling training-free cross-dataset transfer. Extensive experiments demonstrate that the proposed method significantly outperforms baselines in pixel-level average precision across multiple datasets, yielding improvements of 5.63 to 13.23 percentage points while effectively balancing accuracy and efficiency.
📝 Abstract
Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.