LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of CLIP’s vision encoder in producing solely global representations, which hinders its adaptation to region-level tasks. We propose a lightweight adapter-based steering approach that keeps CLIP parameters frozen while employing a MaskAdapter to generate head-wise and layer-wise attention biases injected into self-attention layers, thereby guiding the model to focus on target regions. In the absence of mask inputs, the method seamlessly reverts to the original zero-shot performance. Requiring only minimal additional parameters and small-scale training data, our approach matches or surpasses fully fine-tuned models across multiple benchmarks, achieving a strong synergy between parameter efficiency and feature fidelity.
πŸ“ Abstract
CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation. Our project page is link to https://github.com/AnhKhoa585/lasclip
Problem

Research questions and friction points this paper is trying to address.

CLIP
visual encoder
region-level tasks
global image representations
pre-trained representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight Adapter
Attention Steering
Frozen CLIP
MaskAdapter
Zero-Shot Preservation
πŸ”Ž Similar Papers
No similar papers found.