Score
Design and implement segmentation head modules that convert feature extractor outputs into pixel- or instance-level masks while minimizing parameter count and computational cost; ensure the head is backbone-agnostic and easy to integrate, and tune its architecture and operations to maximize segmentation accuracy under resource constraints.
This work addresses the challenge of fairly evaluating the true performance of diverse backbone architectures in image segmentation, which has been hindered by inconsistencies in decoders, training strategies, and pretraining protocols. To this end, the authors propose LUMA (Lightweight Universal Mask Adapter), a lightweight and architecture-agnostic decoder based on cross-attention, enabling unified training recipes across backbones. Using LUMA, they conduct the first systematic benchmark under identical conditions, evaluating 20 backbones—including ViTs, CNNs, and MoEs—across 11 pretraining strategies and multiple resolutions. Their experiments reveal that the choice of pretraining objective exerts a far greater influence on segmentation performance than architectural design. Moreover, LUMA matches the accuracy of the state-of-the-art efficient ViT-based segmenter EoMT at lower computational cost and demonstrates that so-called “efficient” token mixers offer no advantage at high resolutions, with standard ViTs consistently occupying the Pareto frontier in throughput–accuracy trade-offs.
This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.
Current high-precision interactive segmentation methods face an inherent trade-off between local detail awareness and prompt robustness, limiting their practicality for fine-grained mask generation. This paper introduces the first general-purpose enhancement framework for fine-grained interactive segmentation, preserving SAM2’s generality while overcoming its limitations in local modeling. Our approach comprises three key innovations: (1) localization enhancement—leveraging cross-attention to model local contextual relationships; (2) prompt redirection—spatially aligning and remapping prompt embeddings; and (3) multi-scale mask refinement—employing cascaded feature fusion for progressive mask optimization. Evaluated on both image and video interactive segmentation benchmarks, our method consistently outperforms state-of-the-art approaches, achieving absolute mIoU gains of 3.2–5.7 percentage points. It enables real-time fine-grained editing and temporally consistent cross-frame segmentation, demonstrating significant advances in both accuracy and usability.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
Existing zero-shot image segmentation models (e.g., SAM) yield coarse segmentation masks unsuitable for high-fidelity alpha matte estimation, while conventional matting methods rely on class-specific annotations and thus lack generalizability. To bridge this gap, we propose the first zero-shot image matting framework. Our approach comprises three key components: (1) constructing SA1B-Matte, a large-scale zero-shot matting dataset derived from SA1B; (2) introducing an automatic label conversion pipeline that transforms coarse segmentation masks into pixel-accurate alpha mattes; and (3) designing a hierarchical pixel decoder and prompt-aware mask attention mechanism to enable end-to-end transparency estimation atop the SAM architecture. Evaluated on our newly curated MicroMat-3K benchmark, our method significantly outperforms state-of-the-art approaches. Moreover, it demonstrates strong transferability to downstream tasks including image inpainting and 3D NeRF reconstruction. Code is publicly available.
This work addresses the inefficiency of existing vision backbones on low-parallelism hardware such as CPUs, which are typically optimized for highly parallel accelerators. The authors propose design principles tailored for CPU deployment, emphasizing a balance between high multiply-accumulate operations per second (MACpS) and low latency, and introduce CPUBone—the first family of vision backbones explicitly optimized for CPUs. By incorporating grouped convolutions and small kernel sizes, CPUBone reduces computational load while enhancing execution efficiency on CPU hardware. Experiments demonstrate that CPUBone achieves state-of-the-art accuracy–speed trade-offs across diverse CPU platforms and exhibits strong transfer performance on downstream tasks including object detection and semantic segmentation.
To address the heavy reliance of instance segmentation on costly, labor-intensive manual annotations, this paper proposes a fully unsupervised instance segmentation framework. Methodologically, it introduces the first integration of superpixels—generated via MultiCut and low-level features—with self-supervised visual representations, and designs a superpixel-guided mask loss with dual hard and soft branches. Furthermore, an adaptive-weighted self-training mechanism is incorporated to enable pseudo-label quality-driven iterative optimization. The core contributions are: (1) joint modeling of superpixels and self-supervised features; (2) a two-stage learnable mask loss function; and (3) an adaptive self-training strategy. Evaluated on standard benchmarks, the proposed method achieves state-of-the-art performance in both unsupervised instance segmentation and unsupervised object detection, outperforming all existing approaches.
Existing approaches struggle to balance zero-shot segmentation accuracy with interactive image matting, lacking a unified framework capable of simultaneously achieving high-quality segmentation and fine-grained alpha matte generation. This work proposes SAMA, a lightweight unified model that extends the Segment Anything Model (SAM) by incorporating a Multi-View Local Encoder (MVLE), a Localization Adapter (Local-Adapter), and a dual-task prediction head, thereby integrating interactive segmentation and matting within a single architecture for the first time. Through a joint training strategy, SAMA significantly enhances boundary detail recovery with only a marginal increase in parameters, achieving state-of-the-art performance across multiple segmentation and matting benchmarks and demonstrating its efficiency and versatility for diverse downstream tasks.
This work addresses the high computational cost of conventional fine-tuning for instance segmentation, which typically requires updating a large fraction (40–55%) of parameters in large pre-trained models. To improve parameter efficiency, the study explores parameter-efficient fine-tuning (PEFT) methods, introducing LoRA into deformable attention mechanisms for the first time and systematically evaluating the trade-offs between performance and efficiency based on the number and placement of adapters within the Transformer architecture. Experimental results demonstrate that by fine-tuning only 1–6% of the model parameters, the proposed approach matches or even surpasses the performance of full fine-tuning across four benchmark datasets. These findings validate the efficacy and feasibility of PEFT for instance segmentation and further reveal that its effectiveness is influenced by dataset complexity and model architecture.
该研究针对高分辨率图像语义分割中的效率与准确性平衡问题,提出了一种轻量级框架SiConMo,通过瓶颈阶段的上下文建模来有效整合局部和全局信息。