memory-augmented segmentation

Designs and implements segmentation models and neural networks that incorporate external or learned memory modules (for example, prototype banks, key–value memories, or attention-retrieved exemplars) which are queried to refine or generate pixel-level masks. These systems are built and analyzed to improve robustness to appearance variation, produce consistent masks across images, and to enable segmentation of target regions with limited or no task-specific training by using stored exemplars.

memory-augmentedsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$180K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current high-precision interactive segmentation methods face an inherent trade-off between local detail awareness and prompt robustness, limiting their practicality for fine-grained mask generation. This paper introduces the first general-purpose enhancement framework for fine-grained interactive segmentation, preserving SAM2’s generality while overcoming its limitations in local modeling. Our approach comprises three key innovations: (1) localization enhancement—leveraging cross-attention to model local contextual relationships; (2) prompt redirection—spatially aligning and remapping prompt embeddings; and (3) multi-scale mask refinement—employing cascaded feature fusion for progressive mask optimization. Evaluated on both image and video interactive segmentation benchmarks, our method consistently outperforms state-of-the-art approaches, achieving absolute mIoU gains of 3.2–5.7 percentage points. It enables real-time fine-grained editing and temporally consistent cross-frame segmentation, demonstrating significant advances in both accuracy and usability.

Balancing local detail perception and prompting stabilityEnhancing high-resolution mask generation in foundational segmentation modelsImproving fine-grained segmentation accuracy in images and videos

Demystifying Foreground-Background Memorization in Diffusion Models

Aug 16, 2025
JZ
Jimmy Z. Di
🏛️ University of Waterloo | Vector Institute | University of Ottawa | CISPA Helmholtz Center for Information Security

Diffusion models (DMs) suffer from localized memorization: they not only verbatim reproduce training images but also replicate fine-grained local patches—especially foreground regions—across multiple training samples sharing the same prompt. Existing detection and mitigation methods lack the capability to characterize or suppress such cross-sample, region-level memorization. To address this, we propose FB-Mem—the first foreground-background memory quantification framework grounded in image segmentation. It integrates feature similarity analysis with clustering to separately measure memorization strength in foreground and background regions. Experiments reveal that localized memorization is significantly more pervasive than previously recognized; model-level deactivation techniques prove largely ineffective against foreground memorization; and our clustering-driven data augmentation strategy substantially reduces localized memorization risk. FB-Mem establishes a novel paradigm for memory governance in diffusion models.

Detecting memorization patterns beyond specific prompt-image pairsEvaluating inadequacy of current mitigation methods for local memorizationQuantifying partial memorization in small image regions of diffusion models

RMem: Restricted Memory Banks Improve Video Object Segmentation

Jun 12, 2024
JZ
Junbao Zhou
🏛️ University of Illinois Urbana-Champaign

To address memory redundancy in video object segmentation (VOS)—which impedes feature decoding and causes train-inference misalignment in memory length—this paper proposes a constrained memory mechanism. It restricts memory bank capacity by retaining only semantically critical frames, thereby balancing informativeness and temporal freshness. We are the first to empirically demonstrate that blindly enlarging the memory bank degrades decoder performance. To strengthen inter-frame modeling, we introduce temporal positional encoding. Furthermore, we design a lightweight VOS architecture. Evaluated on VOST (featuring state transitions) and Long Videos benchmarks, our method achieves state-of-the-art performance, significantly improving segmentation accuracy for long videos and dynamic scenes. It also reduces computational overhead and enhances robustness of temporal reasoning.

Long Video ProcessingMemory EfficiencyVideo Object Recognition

On Efficient Variants of Segment Anything Model: A Survey

Oct 07, 2024
XS
Xiaorui Sun
🏛️ UESTC | Lancaster University | Tongji University

While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.

Addressing high computational demands of Segment Anything ModelEnhancing SAM efficiency for resource-limited environmentsSurveying acceleration techniques for SAM variants

Masked Image Modeling: A Survey

Aug 13, 2024
VH
Vlad Hondru
🏛️ University of Bucharest | Amazon | University of Trento

Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.

Automatic LearningComputer VisionMasked Image Modeling

Latest Papers

What's happening recently
View more

Rethinking Memory Design in SAM-Based Visual Object Tracking

Dec 27, 2025
MA
Mohamad Alansari
🏛️ Khalifa University

Current SAM-based visual object tracking models lack principled memory design and exhibit ambiguous cross-generation migration mechanisms. To address this, we propose the first unified hybrid memory framework tailored for SAM architectures, explicitly decoupling short-term appearance memory from long-term distractor suppression memory, and enabling modular integration of diverse memory strategies. Built upon SAM2/SAM3, our framework incorporates object-centric representation, frame-level memory selection, and distractor-aware long-term modeling. We conduct a comprehensive, cross-model evaluation across ten standard benchmarks. Experimental results demonstrate significant improvements in robustness under challenging conditions—including long-term occlusion, complex motion, and strong visual distractors—consistently outperforming SAM2 and SAM3 baselines. The implementation is publicly available.

Analyzes and unifies memory mechanisms across SAM2 and next-generation SAM3 modelsProposes a hybrid memory framework to improve robustness in challenging scenariosSystematically studies memory design in SAM-based visual object tracking

This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.

boundary ambiguitydomain gapfew-shot instance segmentation

Do existing promptable segmentation models genuinely understand semantic concepts, or do they merely rely on visually salient yet semantically misleading cues? This work proposes CAFE, a novel benchmark that systematically evaluates conceptual faithfulness from a counterfactual perspective. By constructing attribute-level counterfactual image pairs—where the target region remains unchanged while misleading appearance, context, or material cues are altered—it assesses model robustness against such distractors. Through text-prompt-guided segmentation evaluation and joint analysis of mask accuracy and semantic consistency across 2,146 samples, the study reveals that models frequently produce high-precision masks in response to incorrect prompts, exposing a significant disconnect between their localization capability and true conceptual understanding.

concept groundingcounterfactual evaluationpromptable segmentation

This work addresses the challenge of amodal instance segmentation in occluded regions, where the absence of pixel observations necessitates reliance on shape priors. The authors propose a reliability-adaptive shape prior framework that dynamically composes instance-specific priors through cross-attention over learnable shape prototypes. To modulate the influence of these priors according to occlusion severity, the method employs the signed distance field of the visible mask as a spatial gating signal, adaptively controlling the strength of prior injection. This approach avoids both uniform prior application and complex generative models, achieving state-of-the-art performance on two standard amodal instance segmentation benchmarks. Under standard evaluation settings, it improves mIoU in occluded regions by over 11 percentage points while using only about one-third the parameters of previous methods.

adaptive injectionamodal instance segmentationinstance mask completion

Hot Scholars

WW

Weiping Wang

School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security
XH

Xiaoshuai Hao

Beijing Academy of Artificial Intelligence,BAAI
vision and language
PJ

Peng Jiang

Kuaishou Technology
Recommender SystemMachine LearningComputational Advertising
XC

Xueqi Cheng

Ph.D. student, Florida State University
Data miningLLMGNNComputational social science
LW

Liyuan Wang

Tsinghua University
bio-inspired learningcontinual learningneuroscience