Score
Designs, builds, and evaluates algorithms and systems that partition input data into coherent segments or labeled regions, producing boundaries, masks, or segment labels and supporting postprocessing such as merging or refinement. Analyzes segmentation accuracy, error modes, and robustness using appropriate evaluation metrics and validation procedures.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
Automatically assessing the success of segmentation refinement is critical for ensuring segmentation reliability and advancing image segmentation techniques. This paper proposes JFS, the first framework to repurpose off-the-shelf few-shot segmentation (FSS) models—such as SegGPT—for segmentation quality assessment. JFS constructs a novel support set from coarse and refined masks and leverages the FSS model to evaluate refinement effectiveness. Crucially, it introduces a mask-driven support set reconstruction mechanism, enabling plug-and-play evaluation without additional training. Experiments on the PASCAL dataset demonstrate that JFS accurately determines the success or failure of mainstream refinement methods (e.g., SEPL), thereby filling a key research gap in automated, reliability-aware refinement assessment. To our knowledge, JFS is the first quantifiable, general-purpose evaluation tool for high-confidence segmentation.
This paper addresses the lack of a unified, comprehensive evaluation framework for image segmentation algorithms. We propose the first holistic, interactive assessment framework encompassing traditional methods (e.g., thresholding, edge detection, region growing), machine learning approaches (e.g., random forests, SVM), and deep learning models (e.g., CNNs, U-Net, Mask R-CNN). Innovatively, we systematically integrate three interaction paradigms—“algorithm-assisted user,” “user-assisted algorithm,” and “hybrid interaction”—and establish a multi-objective evaluation metric balancing segmentation accuracy (IoU), computational efficiency (inference time), and human-factor cost (interaction time). We conduct cross-algorithm benchmarking across diverse scenarios using 12 representative methods, revealing inherent trade-offs among accuracy, speed, and interaction overhead. Experimental results demonstrate that the hybrid interaction paradigm achieves up to a 12.3% IoU improvement in complex scenes, thereby filling a critical gap in the quantitative, comparative evaluation of interactive segmentation systems.
While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.
Random dataset splitting in image segmentation often induces label distribution shifts in the test set, compromising evaluation reliability and model generalization. To address stratification challenges arising from multi-label annotations, strong spatial correlations, and severe class imbalance, we propose two novel stratified splitting methods: Iterative Pixel Stratification (IPS) and Wasserstein-Driven Evolutionary Stratification (WDES). WDES is the first approach to embed the Wasserstein distance into a genetic algorithm framework to globally optimize pixel-level distribution alignment between training and test sets; it further introduces a statistical heterogeneity metric to quantify stratification quality. Extensive validation across diverse domains—including street-scene, medical, and satellite imagery—demonstrates that WDES significantly reduces model performance variance (average reduction of 32.7%) and enhances evaluation robustness under few-shot and long-tailed settings. Our work establishes a verifiable, generalizable theoretical and practical paradigm for data partitioning in semantic segmentation.
This work addresses the limitations of traditional minimal path-based image segmentation methods, which suffer from insufficient accuracy in complex backgrounds and high sensitivity to initialization. To overcome these issues, the authors propose a geodesic framework–based mask proposal voting mechanism that generates diverse and reliable mask candidates through adaptive domain partitioning and constructs the final segmentation via a voting strategy incorporating prior knowledge. The approach innovatively integrates region-based min-cut evolution, geodesic distance computation, and mask scoring maps, effectively eliminating dependence on initial seed points. Experimental results demonstrate that the proposed method significantly outperforms existing minimal path segmentation techniques across multiple challenging scenarios, achieving state-of-the-art performance in both robustness and accuracy.
This work addresses the degradation of model generalization in supervised learning caused by heterogeneity in training data. To mitigate this issue, the authors propose an input-space adaptive partitioning method grounded in the intrinsic heterogeneity of the data. By introducing a variance-based metric that quantifies the inconsistency in pairwise sample influence, they demonstrate that this variance is maximized under mixture distributions. Leveraging this property, the method automatically partitions the data into homogeneous subsets without requiring prior knowledge, enabling independent training of submodels on each subset. Experiments on EMNIST and synthetic datasets show significant improvements in test accuracy, confirming that the proposed variance metric effectively captures data heterogeneity and offers a novel pathway to enhance model generalization.