mask-based detection

Design and implement methods that produce object detections from segmentation masks rather than bounding-box regression, including algorithms to derive bounding boxes from masks, aggregate component masks into instance-level detections, compute mask-based confidence scores, and emit standard detection outputs from segmentation results. This work covers the modeling, post-processing, and evaluation procedures needed to convert, score, and merge masks into detection-format outputs.

mask-baseddetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Extending Dataset Pruning to Object Detection: A Variance-based Approach

May 22, 2025
RY
Ryota Yagi
🏛️ University of Nevada, Reno

This work pioneers the extension of dataset pruning to object detection, addressing three core challenges: absence of object-level attribution, lack of detection-specific scoring mechanisms, and difficulty in aggregating image-level sample quality. We propose Variance-driven Prediction Scoring (VPS), which jointly leverages IoU and confidence scores for object-level sample assessment. Furthermore, we establish the first theoretical framework for object detection dataset pruning, incorporating multi-scale feature response aggregation and confidence-weighted quality fusion. Extensive experiments on PASCAL VOC and MS COCO demonstrate consistent mAP improvements of 2.1–3.8 percentage points over state-of-the-art methods. Empirical analysis reveals that sample informativeness—rather than sheer data volume or class balance—yields greater training efficiency and model performance gains.

Addressing key challenges in pruning for detectionExtending dataset pruning to object detection tasksProposing Variance-based Prediction Score for sample selection

This work proposes a novel turbo inference strategy that enables bidirectional iterative refinement between detection and segmentation during inference, without requiring retraining. Challenging the conventional unidirectional “detect-then-segment” paradigm in top-down instance segmentation, the method introduces a closed-loop interaction mechanism through turbo detection and segmentation heads, coupled with cross-task feature fusion. This design dynamically leverages complementary information from both tasks while preserving the top-down architecture. Extensive experiments demonstrate significant improvements in both detection and segmentation accuracy on COCO, iFLYTEK, and Cityscapes benchmarks, achieving a favorable trade-off between performance and inference efficiency.

detect-then-segment paradigminstance segmentationobject detection

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

Confidence-driven Bounding Box Localization for Small Object Detection

Mar 03, 2023
HS
Huixin Sun
🏛️ Beihang University

Small-object detection suffers from significant performance degradation due to inaccurate localization and unstable gradients. This paper identifies that conventional regression-based bounding box localization induces distorted gradients for small objects. To address this, we reformulate bounding box localization as a grid-based classification task—the first such approach—and propose a confidence-driven localization framework. Our method employs two-hot label encoding, confidence distribution prediction, cross-entropy localization loss, and an entropy-based uncertainty loss to jointly correct gradient flow and suppress localization uncertainty. Evaluated on mainstream detectors—including YOLOv8 and RT-DETR—and across three major benchmarks (COCO, VisDrone, and AI-TOD), our approach achieves state-of-the-art performance, notably improving AP for small objects. Moreover, it demonstrates strong generalization across diverse annotation protocols and high-resolution imagery.

Addresses gradient instability in small object localizationImproves small object detection performance gapProposes uncertainty-aware framework for stable gradients

Existing interpretability methods for object detection predominantly rely on single-pixel attribution, failing to capture the joint influence of multi-pixel collaborations on both bounding box localization and class prediction—thus overlooking compositional cues or introducing spurious correlations. To address this, we introduce Shapley interaction values to object detection interpretation for the first time, proposing the first end-to-end differentiable framework that explicitly models high-order cooperative effects among pixel groups. Our approach jointly models feature-space perturbations and detection output sensitivity to simultaneously quantify individual pixel contributions and higher-order interactions. Extensive experiments on COCO and other benchmarks demonstrate that our method significantly outperforms state-of-the-art interpretability baselines. Both qualitative visualizations and quantitative metrics confirm its superior ability to localize discriminative visual regions accurately. The source code will be made publicly available.

Captures both individual and collective pixel influencesExplains object detectors via collective pixel contributionsProvides explanations for bounding box and class decisions

Latest Papers

What's happening recently
View more

Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.

binary segmentationevaluation metricsmetric decomposition

This work addresses the challenges of ambiguous boundaries and missed detections in weakly supervised camouflaged object detection caused by coarse bounding box annotations. To tackle these issues, the authors propose MGNet, a novel framework that innovatively integrates the Segment Anything Model (SAM) with weakly supervised learning through a BoxSAM strategy to generate high-quality pixel-level pseudo-labels. The architecture further incorporates a cascaded mask decoder, a Context Enhancement Module (CEM), and a Mask-Guided Feature Aggregation Module (MFAM) to enable mask-guided fine-grained segmentation. Experimental results demonstrate that MGNet significantly outperforms existing weakly supervised methods across multiple benchmarks, achieving performance on par with current state-of-the-art approaches.

Annotation EfficiencyCamouflaged Object DetectionEdge Ambiguity

This work addresses the scarcity of pixel-level annotations in industrial inspection and the systematic noise in pseudo-masks generated by Segment Anything Model (SAM) from bounding boxes, which often leads to false positives on background regions or missed sparse defects. To tackle this, the authors propose a noise-robust box-to-pixel distillation framework that treats SAM as a noisy teacher model to generate offline pseudo-masks and trains a lightweight student model for weakly supervised defect segmentation. The approach incorporates a hierarchical decoder with an auxiliary binary localization head to decouple foreground discovery from classification and introduces a unidirectional online self-correction mechanism to mitigate the teacher’s false negatives. Evaluated on a wind turbine inspection benchmark, the method achieves significant gains: +6.97 in anomaly mIoU, +9.71 in binary IoU, and +18.56 in recall, while reducing trainable parameters by 80%.

defect segmentationindustrial inspectionnoisy pseudo-labels

This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.

boundary ambiguitydomain gapfew-shot instance segmentation

This study addresses the challenge of effectively leveraging large-scale unlabeled images to improve object detection performance under limited annotation budgets. We systematically evaluate three prominent semi-supervised object detection methods—MixPL, Semi-DETR, and Consistent-Teacher—across MS-COCO, Pascal VOC, and a custom Beetle dataset, analyzing their trade-offs among accuracy, model size, and inference latency under varying labeling ratios. For the first time, we reveal consistent patterns of performance degradation as labeled data decreases, both on general-purpose and domain-specific datasets. Our empirical findings provide actionable insights and practical guidance for selecting appropriate semi-supervised approaches in resource-constrained scenarios.

data-scarce learningfew-shot learningobject detection

Hot Scholars

BX

Bin Xie

InfoBeyond Technology LLC
Mobile ComuptingSecurityBig Data Streaming
KZ

Kun Zhan

AI Researcher, LiAuto
Autonomous DrivingComputer Vision3D Vision
LJ

Licheng Jiao

Distinguished Professor of Xidian University, IEEE Fellow
Neural NetworksComputational IntelligenceEvolutionary ComputationRemote Sensing
ES

Eric Sax

Karlsruhe Institute for Technology
Systems Engineering
JL

Jiasen Lu

Research Scientist, Apple
Computer VisionNatural Language Processing