Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of visual matching failures in few-shot segmentation caused by appearance variations and occlusions. To this end, it proposes the first few-shot segmentation framework that leverages the reasoning capabilities of multimodal large language models (MLLMs). Built upon SAM 2, the framework exploits MLLMs to extract spatial and semantic priors. Furthermore, a dual-memory debate fusion mechanism is designed to enhance the robustness of target representations, alongside a progressive cross-modal prompt generation module that facilitates efficient synergy among multimodal features. Extensive experiments across multiple benchmarks demonstrate that the proposed approach significantly improves segmentation accuracy in complex scenarios, achieving substantial performance gains over existing state-of-the-art methods.
📝 Abstract
Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.
Problem

Research questions and friction points this paper is trying to address.

Few-Shot Segmentation
Target Knowledge
Appearance Discrepancies
Occlusions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Shot Segmentation
Multimodal Large Language Models
Dual-Memory Debate-Fusion
Progressive Cross-Modal Prompt Generator
SAM 2
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.