PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that small-object features in vision-language encoders are prone to confusion with background, hindering accurate localization. To this end, it proposes a prototype-contrastive and local-magnification mechanism built upon a frozen CLIP backbone. Foreground–background prototypes are constructed from the support set, and precise localization is achieved through densely sampled query windows, reprojection, and coverage averaging. The per-class memory footprint remains fixed, enabling training-free cross-dataset transfer. Extensive experiments demonstrate that the proposed method significantly outperforms baselines in pixel-level average precision across multiple datasets, yielding improvements of 5.63 to 13.23 percentage points while effectively balancing accuracy and efficiency.
📝 Abstract
Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.
Problem

Research questions and friction points this paper is trying to address.

small-target localization
frozen CLIP
vision-language encoder
spatial feature mixing
Innovation

Methods, ideas, or system contributions that make the work stand out.

frozen CLIP
prototype contrast
local magnification
small-target localization
support-conditioned
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhipeng Ye
Taizhou Institute of Science and Technology, Nanjing University of Science and Technology, Taizhou 225300, Jiangsu, China
Feng Jiang
Feng Jiang
Shenzhen University of Advanced Technology
Discourse ParsingLarge-scale Language ModelDialogue System
Q
Qiufeng Wang
Department of Intelligence Science, Xi’an Jiaotong-Liverpool University, Suzhou 215123, Jiangsu, China
H
Hao Li
School of Computer Science and Technology, University of Arizona, Tucson 85705, AZ, USA