single-stage weakly supervised segmentation

Designs, builds, and evaluates end-to-end single-stage segmentation models that learn pixel- or region-level masks from weak supervision (e.g., image-level, scribble, or bounding-box labels) without relying on an offline pseudo-mask refinement stage. Focuses on architectures, loss functions, and training procedures that generate activation maps and final masks in one pass to reduce training time and prevent error propagation across stages.

single-stageweaklysupervisedsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high cost of dense pixel-wise annotations in semantic segmentation by proposing SeSAM, a framework that systematically analyzes and mitigates the limitations of the Segment Anything Model (SAM) under weak supervision. SeSAM effectively transfers SAM’s instance-level priors to semantic segmentation through a pipeline comprising category-aware mask decomposition, skeleton-point-based prompt sampling, weak-label-guided mask selection, and iterative pseudo-label refinement. By leveraging coarse masks, scribbles, or point annotations as weak supervisory signals, the method achieves substantial performance gains across diverse weakly supervised settings. Experimental results demonstrate that SeSAM significantly outperforms existing approaches while approaching the accuracy of fully supervised models, thereby offering a practical pathway to drastically reduce annotation costs without compromising segmentation quality.

annotation costclass-based segmentationinstance priors

Current high-precision interactive segmentation methods face an inherent trade-off between local detail awareness and prompt robustness, limiting their practicality for fine-grained mask generation. This paper introduces the first general-purpose enhancement framework for fine-grained interactive segmentation, preserving SAM2’s generality while overcoming its limitations in local modeling. Our approach comprises three key innovations: (1) localization enhancement—leveraging cross-attention to model local contextual relationships; (2) prompt redirection—spatially aligning and remapping prompt embeddings; and (3) multi-scale mask refinement—employing cascaded feature fusion for progressive mask optimization. Evaluated on both image and video interactive segmentation benchmarks, our method consistently outperforms state-of-the-art approaches, achieving absolute mIoU gains of 3.2–5.7 percentage points. It enables real-time fine-grained editing and temporally consistent cross-frame segmentation, demonstrating significant advances in both accuracy and usability.

Balancing local detail perception and prompting stabilityEnhancing high-resolution mask generation in foundational segmentation modelsImproving fine-grained segmentation accuracy in images and videos

This study addresses the difficulty of disentangling structural contributions from model capacity in mask refinement, as well as its poor cross-generator generalization. We formulate the refinement process as conditional random field (CRF) inference. Specifically, a zero-state mechanism is introduced to handle weakly labeled regions, decoupling boundary from non-boundary errors. Furthermore, DINOv2 features are leveraged to construct image-conditioned latent region consistency constraints, and damped mean-field iterations are unrolled to perform structured corrections. This approach establishes a verification paradigm demonstrating that explicit structure outperforms black-box capacity. Extensive experiments on datasets such as COD10K show that our method significantly surpasses parameter-matched baselines, achieves performance comparable to foundation models with minimal parameters, and effectively enhances generalization to unseen generators.

conditional random fieldgeneralizationimage segmentation

From Few to More: Scribble-based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo Labels

Aug 23, 2024
ZW
Zhisong Wang
🏛️ Northwestern Polytechnical University | Shandong Artificial Intelligence Institute | Qilu University of Technology (Shandong Academy of Sciences) | Ningbo Institute of Northwestern Polytechnical University | Research & Development Institute of Northwestern Polytechnical University in Shenzhen

To address performance degradation in weakly supervised medical image segmentation caused by sparse scribble annotations, this paper proposes MaCo, a “from-few-to-many” progressive learning framework. Methodologically, we introduce Mask Context Modeling (MCM), a novel attention mechanism that jointly leverages distance maps and an exponential decay function to generate continuous pseudo-labels (CPLs) with pixel-wise semantic confidence—replacing error-prone hard pseudo-labels. Additionally, we incorporate a self-supervised context consistency constraint, eliminating reliance on auxiliary tasks. This enables robust, continuous modeling of pixel-level semantic confidence. Evaluated on three public medical imaging benchmarks, MaCo consistently outperforms existing scribble-supervised methods, establishing new state-of-the-art results and significantly enhancing model adaptability to annotation sparsity.

Addresses sparse annotation challenges in segmentation modelsEnhances semantic consistency without hard pseudo labelsImproves scribble-based medical image segmentation accuracy

On Efficient Variants of Segment Anything Model: A Survey

Oct 07, 2024
XS
Xiaorui Sun
🏛️ UESTC | Lancaster University | Tongji University

While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.

Addressing high computational demands of Segment Anything ModelEnhancing SAM efficiency for resource-limited environmentsSurveying acceleration techniques for SAM variants

Latest Papers

What's happening recently
View more

This work addresses the challenges of low-quality pseudo-labels and difficulty in incorporating domain priors in weakly supervised semantic segmentation by proposing a neuro-symbolic approach that, for the first time, integrates differentiable fuzzy logic as a semantic regularization mechanism during foundation model fine-tuning. The method unifies diverse weak annotations—such as image-level tags and bounding boxes—with domain knowledge into continuous logical constraints, which are then used to refine pseudo-labels generated by the Segment Anything Model (SAM) and subsequently train a prompt-free segmentation model. Evaluated on Pascal VOC 2012 and REFUGE2, the approach substantially outperforms existing weakly supervised methods and even surpasses most fully supervised baselines, demonstrating its effectiveness and superiority in jointly modeling heterogeneous weak supervision signals and prior knowledge.

Foundation ModelsHeterogeneous LabelsPrior Knowledge

This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.

boundary ambiguitydomain gapfew-shot instance segmentation

This work addresses the scarcity of high-quality mask–text pairs and the optimization conflict between region segmentation and understanding tasks in pixel-level multimodal large language models, which arises from disparities in supervision formats and signal density. To overcome these challenges, the authors propose PixVL, a self-supervised post-training framework that enables the model to generate and self-verify region descriptions using unlabeled data through a mask–text consistency loop. Key innovations include confusion-aware semantic validation, cross-view consistency checks (e.g., across video frames or geometric transformations), and a quality-coupled bidirectional learning strategy that formulates segmentation and understanding as a synergistic generator–verifier mechanism. Experiments demonstrate that PixVL significantly improves performance on both tasks, effectively mitigating optimization interference while efficiently leveraging unlabeled data.

mask-text pairsoptimization interferencepixel-level MLLMs

This study addresses the issues of error fixation caused by hard labels and insufficient boundary precision in weakly supervised semantic segmentation by proposing the DS-CRF framework. Inspired by Conditional Random Fields, this method decouples unary and pairwise potential supervision signals. Specifically, it leverages DINO text CAMs to provide soft classification supervision and incorporates SAM to extract boundary information, optimizing training via a custom CRF loss function. This design effectively prevents the error amplification typically induced by conventional pseudo-label fusion while preserving prediction uncertainty to enhance boundary learning. Experimental results demonstrate that the proposed framework achieves 56.5% mIoU on the MS COCO dataset, establishing a new state-of-the-art performance.

Boundary RefinementClass Activation MapsDenseCRF

This work addresses common limitations in image segmentation models—such as ambiguous boundaries, semantic inconsistency, and structural errors—by introducing the Phoenix framework. Phoenix generates semantically aware noise through adversarial mask perturbations to simulate realistic segmentation errors and employs a contrastive learning–based tripartite refinement mechanism that simultaneously enhances intra-class feature consistency and inter-class separability. Integrating adversarial learning, embedding attacks, and relational modeling, Phoenix operates as a plug-and-play module without requiring modifications to the backbone architecture. Extensive experiments demonstrate that Phoenix consistently outperforms existing approaches across diverse segmentation tasks, delivering substantial improvements in mask quality and reliably boosting the performance of state-of-the-art models.

adversarial perturbationboundary imperfectionmask refinement

Hot Scholars

XH

Xiaowei Huang

Professor of Computer Science, University of Liverpool
AI Safety and SecurityVerificationTrustworthy AIFormal Methods
JX

Jimin Xiao

Professor in Intelligent Science, Xi'an Jiaotong-Liverpool University
computer visionmachine learning
CK

Christopher Kanan

University of Rochester
Artificial IntelligenceDeep LearningAGIMulti-Modal AI