training-free few-shot segmentation

Design and implement segmentation systems that, at inference time, assign pixel-level labels to a query image using only a small set of labeled support examples, without any additional training or parameter adaptation. These systems build inference-time modules that leverage frozen pretrained feature extractors or semantic priors, similarity-based label propagation or matching, and mechanisms to aggregate or sharpen decisions as more references accumulate.

training-freefew-shotsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Reasoning Segmentation for Images and Videos: A Survey

May 24, 2025
YS
Yiqing Shen
🏛️ Johns Hopkins University | Amazon Web Services

This paper addresses the fundamental gap between visual perception and human-like reasoning by proposing “Reasoning Segmentation” (RS)—a novel paradigm that achieves precise open-vocabulary segmentation of objects in images/videos via implicit textual queries, integrated commonsense knowledge, and logical inference. We establish the first comprehensive survey framework for RS, formally defining its three core characteristics: language-driven segmentation, knowledge-enhanced representation, and dynamic reasoning. Our systematic review encompasses 26 methods, 29 benchmark datasets, and standardized evaluation protocols. Technically, we unify multimodal large models, vision-language alignment, prompt engineering, and knowledge injection into a coherent taxonomic structure. Through empirical analysis, we identify key performance bottlenecks and propose six promising future directions: scalable reasoning architectures, embodied RS, causal segmentation, neuro-symbolic integration, interactive RS, and foundation-model-based RS. This work lays the conceptual and methodological groundwork for advancing reasoning-aware visual understanding.

Bridge visual perception and human-like reasoning via natural languageDelineate objects using implicit text queries requiring reasoningSurvey reasoning segmentation methods, metrics, datasets, and applications

Few-Shot Segmentation with Global and Local Contrastive Learning

Aug 11, 2021
WL
Weide Liu
🏛️ Nanyang Technological University | Fudan University | Institute for Infocomm Research

This work addresses the over-reliance on support images and the difficulty in modeling query-specific discriminative priors in few-shot segmentation. We propose a **query self-driven prior extraction paradigm**, the first to explicitly eliminate dependence on support images and instead learn object localization priors unsupervised solely from the query image. Methodologically, we design a global-local contrastive learning framework to train a lightweight prior extractor that generates prior region maps and guides cross-branch feature interaction. This decouples query feature learning from support sample binding, substantially improving generalization. Our approach achieves new state-of-the-art performance on PASCAL-5^i and COCO, delivering significant gains without resorting to complex modules. It offers a more interpretable, low-coupling modeling perspective for few-shot segmentation.

Extracting query information independently for few-shot segmentationImproving segmentation via global-local contrastive learningOvercoming support dependency in query feature extraction

Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts

Jul 02, 2024
PD
Pasquale De Marinis
🏛️ University of Bari Aldo Moro | NVIDIA AI Technology Center

Existing multi-class few-shot semantic segmentation methods suffer from weak generalization, limited prompt modality support (e.g., only points or boxes), and strong dependence on fixed N-way K-shot configurations. Method: We propose the first unified end-to-end framework for multi-class few-shot segmentation, built upon a vision-language promptable Transformer architecture that accepts arbitrary combinations of point, box, and mask prompts. It introduces multi-prompt fusion encoding and cross-class contrastive learning to construct a class-agnostic, unified support-set representation. Contribution/Results: Our key innovation is the complete decoupling of the model from both the number of support classes (N) and the number of shots per class (K), enabling truly generalizable few-shot segmentation. Evaluated on COCO-20i, our method achieves state-of-the-art performance while significantly reducing training overhead and maintaining robustness across all settings—from 1-way 1-shot to arbitrary N-way K-shot configurations.

Achieves state-of-the-art generalization without retrainingEnables multi-class few-shot segmentation with minimal examplesIntroduces versatile visual prompts for enhanced adaptability

This work addresses the limited interpretability of semantic segmentation models by proposing ProtoSeg, an interpretable segmentation framework based on prototype retrieval. Methodologically, ProtoSeg models each semantic class as a set of visual prototypes and performs patch-level similarity matching against the training set to retrieve the most relevant image patches for prediction support. A diversity loss is introduced to encourage prototypes to capture semantically rich, discriminative, and intra-class-comprehensive local parts. ProtoSeg establishes, for the first time, an “prototype–part–semantics” aligned interpretable segmentation paradigm. Evaluated on Pascal VOC and Cityscapes, ProtoSeg achieves segmentation accuracy competitive with strong baselines while providing intuitive, verifiable decision rationales—significantly enhancing model transparency and human interpretability.

Ensures accuracy comparable to baseline methodsImproves prototype diversity within classesInterpretable semantic image segmentation model

Group-On: Boosting One-Shot Segmentation with Supportive Query

Apr 18, 2024
HZ
Hanjing Zhou
🏛️ Zhejiang University | University of Illinois at Urbana-Champaign | University of Notre Dame

One-shot semantic segmentation suffers from performance degradation due to large intra- and inter-class variations in appearance and pose; existing methods rely on multiple annotated support images, incurring high labeling costs. This paper proposes a “query-as-support” co-optimization paradigm: within a batch, query images serve as mutual pseudo-support samples—eliminating the need for additional annotations. We design a lightweight Group-On Voting module to enable cross-query mask mutual enhancement. Built upon ASNet/HSNet, our framework establishes intra-batch query interaction through coarse prediction, pseudo-support construction, and voting-based fusion. On COCO-20i, our method achieves mIoU improvements of 8.21% and 7.46% over ASNet and HSNet, respectively. Notably, its one-shot performance surpasses most five-shot approaches, marking the first time that one-shot segmentation accuracy approaches that of multi-shot settings.

Addresses one-shot semantic segmentation with intra-class variation challengesEnhances segmentation accuracy via mutual query-mask supportReduces manual labeling costs by using pseudo support data

Latest Papers

What's happening recently
View more

This study addresses the challenge of spatial misalignment between remote sensing imagery and external labels—such as those from OpenStreetMap—which hinders effective training of building segmentation models. To overcome this, the authors propose Align and Segment (AnS), an unsupervised framework that jointly performs high-quality building segmentation and automatic image-to-label alignment without requiring precisely registered annotations. The method employs a differentiable spatial transformation module to affinely align labels and incorporates a self-supervised regularization loss to prevent shortcut learning. This work is the first to simultaneously achieve accurate segmentation and alignment under misaligned supervision, thereby relaxing the conventional reliance on pixel-perfect ground truth. Extensive experiments on real and synthetic datasets across multiple cities demonstrate the superiority of AnS in both segmentation accuracy and alignment fidelity.

building segmentationimage-label alignmentmisaligned labels

Do existing promptable segmentation models genuinely understand semantic concepts, or do they merely rely on visually salient yet semantically misleading cues? This work proposes CAFE, a novel benchmark that systematically evaluates conceptual faithfulness from a counterfactual perspective. By constructing attribute-level counterfactual image pairs—where the target region remains unchanged while misleading appearance, context, or material cues are altered—it assesses model robustness against such distractors. Through text-prompt-guided segmentation evaluation and joint analysis of mask accuracy and semantic consistency across 2,146 samples, the study reveals that models frequently produce high-precision masks in response to incorrect prompts, exposing a significant disconnect between their localization capability and true conceptual understanding.

concept groundingcounterfactual evaluationpromptable segmentation

This work addresses the instability and lack of robustness in existing in-context segmentation methods, which produce inconsistent results for the same query image under varying reference images. To tackle this issue, the study reformulates the task from a robustness perspective and introduces a concept-guided in-context segmentation paradigm. It leverages a multimodal large language model to generate high-level semantic concepts, which—combined with visual exemplars—jointly activate a frozen SAM3 model. A concept reasoning module and a tree-search optimization mechanism are further integrated to enable synergistic semantic guidance and spatial localization. The proposed approach achieves state-of-the-art accuracy on standard benchmarks while significantly reducing output variance across different reference images, thereby substantially enhancing system robustness.

in-context segmentationreference variabilityrobustness

This work addresses the limited generalization of existing promptable segmentation methods on context-dependent and reasoning-intensive complex concepts. It formalizes concept segmentation as a rule-guided concept grounding task and introduces a three-tier conceptual taxonomy—comprising Concrete Instance (CI), Contextual Description (CD), and Compositional Reasoning (CR)—to characterize cognitive complexity. The authors propose Meta-GRPO, a meta-reinforcement learning mechanism that learns transferable rules from visual exemplars and integrates agent-based reasoning with a lightweight translation module to enable end-to-end generation of segmentation prompts from reasoning states, while preserving an efficient inference path for simple scenarios. Compatible with mainstream promptable segmentation backbones, the method demonstrates comprehensive coverage across all conceptual tiers on CR benchmarks spanning natural, industrial, and medical domains, significantly improving performance on complex concept segmentation without compromising baseline efficiency.

cognitive complexityconcept segmentationgeneralization

This work addresses the challenges of low-quality pseudo-labels and difficulty in incorporating domain priors in weakly supervised semantic segmentation by proposing a neuro-symbolic approach that, for the first time, integrates differentiable fuzzy logic as a semantic regularization mechanism during foundation model fine-tuning. The method unifies diverse weak annotations—such as image-level tags and bounding boxes—with domain knowledge into continuous logical constraints, which are then used to refine pseudo-labels generated by the Segment Anything Model (SAM) and subsequently train a prompt-free segmentation model. Evaluated on Pascal VOC 2012 and REFUGE2, the approach substantially outperforms existing weakly supervised methods and even surpasses most fully supervised baselines, demonstrating its effectiveness and superiority in jointly modeling heterogeneous weak supervision signals and prior knowledge.

Foundation ModelsHeterogeneous LabelsPrior Knowledge

Hot Scholars

MR

Mirabela Rusu

Assistant Professor of Radiology at Stanford University
multi-protocolmulti-scale data fusionMRIHistology
NK

Neelesh Kumar

Senior Scientist - AI R&D, Procter and Gamble
Machine Learning
AS

Assaf Shocher

Assistant Professor at Technion
Deep LearningComputer Vision
SU

Shaheer U. Saeed

University College London
Machine LearningMedical Image ComputingReinforcement Learning