Score
Designs and implements segmentation models and neural networks that incorporate external or learned memory modules (for example, prototype banks, key–value memories, or attention-retrieved exemplars) which are queried to refine or generate pixel-level masks. These systems are built and analyzed to improve robustness to appearance variation, produce consistent masks across images, and to enable segmentation of target regions with limited or no task-specific training by using stored exemplars.
Current high-precision interactive segmentation methods face an inherent trade-off between local detail awareness and prompt robustness, limiting their practicality for fine-grained mask generation. This paper introduces the first general-purpose enhancement framework for fine-grained interactive segmentation, preserving SAM2’s generality while overcoming its limitations in local modeling. Our approach comprises three key innovations: (1) localization enhancement—leveraging cross-attention to model local contextual relationships; (2) prompt redirection—spatially aligning and remapping prompt embeddings; and (3) multi-scale mask refinement—employing cascaded feature fusion for progressive mask optimization. Evaluated on both image and video interactive segmentation benchmarks, our method consistently outperforms state-of-the-art approaches, achieving absolute mIoU gains of 3.2–5.7 percentage points. It enables real-time fine-grained editing and temporally consistent cross-frame segmentation, demonstrating significant advances in both accuracy and usability.
Diffusion models (DMs) suffer from localized memorization: they not only verbatim reproduce training images but also replicate fine-grained local patches—especially foreground regions—across multiple training samples sharing the same prompt. Existing detection and mitigation methods lack the capability to characterize or suppress such cross-sample, region-level memorization. To address this, we propose FB-Mem—the first foreground-background memory quantification framework grounded in image segmentation. It integrates feature similarity analysis with clustering to separately measure memorization strength in foreground and background regions. Experiments reveal that localized memorization is significantly more pervasive than previously recognized; model-level deactivation techniques prove largely ineffective against foreground memorization; and our clustering-driven data augmentation strategy substantially reduces localized memorization risk. FB-Mem establishes a novel paradigm for memory governance in diffusion models.
To address memory redundancy in video object segmentation (VOS)—which impedes feature decoding and causes train-inference misalignment in memory length—this paper proposes a constrained memory mechanism. It restricts memory bank capacity by retaining only semantically critical frames, thereby balancing informativeness and temporal freshness. We are the first to empirically demonstrate that blindly enlarging the memory bank degrades decoder performance. To strengthen inter-frame modeling, we introduce temporal positional encoding. Furthermore, we design a lightweight VOS architecture. Evaluated on VOST (featuring state transitions) and Long Videos benchmarks, our method achieves state-of-the-art performance, significantly improving segmentation accuracy for long videos and dynamic scenes. It also reduces computational overhead and enhances robustness of temporal reasoning.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
Current SAM-based visual object tracking models lack principled memory design and exhibit ambiguous cross-generation migration mechanisms. To address this, we propose the first unified hybrid memory framework tailored for SAM architectures, explicitly decoupling short-term appearance memory from long-term distractor suppression memory, and enabling modular integration of diverse memory strategies. Built upon SAM2/SAM3, our framework incorporates object-centric representation, frame-level memory selection, and distractor-aware long-term modeling. We conduct a comprehensive, cross-model evaluation across ten standard benchmarks. Experimental results demonstrate significant improvements in robustness under challenging conditions—including long-term occlusion, complex motion, and strong visual distractors—consistently outperforming SAM2 and SAM3 baselines. The implementation is publicly available.
This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.
Do existing promptable segmentation models genuinely understand semantic concepts, or do they merely rely on visually salient yet semantically misleading cues? This work proposes CAFE, a novel benchmark that systematically evaluates conceptual faithfulness from a counterfactual perspective. By constructing attribute-level counterfactual image pairs—where the target region remains unchanged while misleading appearance, context, or material cues are altered—it assesses model robustness against such distractors. Through text-prompt-guided segmentation evaluation and joint analysis of mask accuracy and semantic consistency across 2,146 samples, the study reveals that models frequently produce high-precision masks in response to incorrect prompts, exposing a significant disconnect between their localization capability and true conceptual understanding.
本文针对复杂时间动态下的视频对象分割问题,提出了一种竞争性记忆读取方法,以提高目标身份保持的准确性。
This work addresses the challenge of amodal instance segmentation in occluded regions, where the absence of pixel observations necessitates reliance on shape priors. The authors propose a reliability-adaptive shape prior framework that dynamically composes instance-specific priors through cross-attention over learnable shape prototypes. To modulate the influence of these priors according to occlusion severity, the method employs the signed distance field of the visible mask as a spatial gating signal, adaptively controlling the strength of prior injection. This approach avoids both uniform prior application and complex generative models, achieving state-of-the-art performance on two standard amodal instance segmentation benchmarks. Under standard evaluation settings, it improves mIoU in occluded regions by over 11 percentage points while using only about one-third the parameters of previous methods.