Score
Design and implement methods that produce object detections from segmentation masks rather than bounding-box regression, including algorithms to derive bounding boxes from masks, aggregate component masks into instance-level detections, compute mask-based confidence scores, and emit standard detection outputs from segmentation results. This work covers the modeling, post-processing, and evaluation procedures needed to convert, score, and merge masks into detection-format outputs.
This work pioneers the extension of dataset pruning to object detection, addressing three core challenges: absence of object-level attribution, lack of detection-specific scoring mechanisms, and difficulty in aggregating image-level sample quality. We propose Variance-driven Prediction Scoring (VPS), which jointly leverages IoU and confidence scores for object-level sample assessment. Furthermore, we establish the first theoretical framework for object detection dataset pruning, incorporating multi-scale feature response aggregation and confidence-weighted quality fusion. Extensive experiments on PASCAL VOC and MS COCO demonstrate consistent mAP improvements of 2.1–3.8 percentage points over state-of-the-art methods. Empirical analysis reveals that sample informativeness—rather than sheer data volume or class balance—yields greater training efficiency and model performance gains.
This work proposes a novel turbo inference strategy that enables bidirectional iterative refinement between detection and segmentation during inference, without requiring retraining. Challenging the conventional unidirectional “detect-then-segment” paradigm in top-down instance segmentation, the method introduces a closed-loop interaction mechanism through turbo detection and segmentation heads, coupled with cross-task feature fusion. This design dynamically leverages complementary information from both tasks while preserving the top-down architecture. Extensive experiments demonstrate significant improvements in both detection and segmentation accuracy on COCO, iFLYTEK, and Cityscapes benchmarks, achieving a favorable trade-off between performance and inference efficiency.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
Small-object detection suffers from significant performance degradation due to inaccurate localization and unstable gradients. This paper identifies that conventional regression-based bounding box localization induces distorted gradients for small objects. To address this, we reformulate bounding box localization as a grid-based classification task—the first such approach—and propose a confidence-driven localization framework. Our method employs two-hot label encoding, confidence distribution prediction, cross-entropy localization loss, and an entropy-based uncertainty loss to jointly correct gradient flow and suppress localization uncertainty. Evaluated on mainstream detectors—including YOLOv8 and RT-DETR—and across three major benchmarks (COCO, VisDrone, and AI-TOD), our approach achieves state-of-the-art performance, notably improving AP for small objects. Moreover, it demonstrates strong generalization across diverse annotation protocols and high-resolution imagery.
Existing interpretability methods for object detection predominantly rely on single-pixel attribution, failing to capture the joint influence of multi-pixel collaborations on both bounding box localization and class prediction—thus overlooking compositional cues or introducing spurious correlations. To address this, we introduce Shapley interaction values to object detection interpretation for the first time, proposing the first end-to-end differentiable framework that explicitly models high-order cooperative effects among pixel groups. Our approach jointly models feature-space perturbations and detection output sensitivity to simultaneously quantify individual pixel contributions and higher-order interactions. Extensive experiments on COCO and other benchmarks demonstrate that our method significantly outperforms state-of-the-art interpretability baselines. Both qualitative visualizations and quantitative metrics confirm its superior ability to localize discriminative visual regions accurately. The source code will be made publicly available.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This work addresses the challenges of ambiguous boundaries and missed detections in weakly supervised camouflaged object detection caused by coarse bounding box annotations. To tackle these issues, the authors propose MGNet, a novel framework that innovatively integrates the Segment Anything Model (SAM) with weakly supervised learning through a BoxSAM strategy to generate high-quality pixel-level pseudo-labels. The architecture further incorporates a cascaded mask decoder, a Context Enhancement Module (CEM), and a Mask-Guided Feature Aggregation Module (MFAM) to enable mask-guided fine-grained segmentation. Experimental results demonstrate that MGNet significantly outperforms existing weakly supervised methods across multiple benchmarks, achieving performance on par with current state-of-the-art approaches.
This work addresses the scarcity of pixel-level annotations in industrial inspection and the systematic noise in pseudo-masks generated by Segment Anything Model (SAM) from bounding boxes, which often leads to false positives on background regions or missed sparse defects. To tackle this, the authors propose a noise-robust box-to-pixel distillation framework that treats SAM as a noisy teacher model to generate offline pseudo-masks and trains a lightweight student model for weakly supervised defect segmentation. The approach incorporates a hierarchical decoder with an auxiliary binary localization head to decouple foreground discovery from classification and introduces a unidirectional online self-correction mechanism to mitigate the teacher’s false negatives. Evaluated on a wind turbine inspection benchmark, the method achieves significant gains: +6.97 in anomaly mIoU, +9.71 in binary IoU, and +18.56 in recall, while reducing trainable parameters by 80%.
This work addresses the challenge of instance segmentation in low-texture, uniformly colored industrial scenes, where existing methods struggle due to insufficient texture and contextual cues, and where instance definitions often vary by task. The authors propose a boundary-aware few-shot instance segmentation framework that leverages a pretrained foundation model to extract visual features and employs a lightweight signed distance function (SDF) head to predict boundary-sensitive distance maps. Instance masks are then reconstructed from these SDF predictions. By shifting supervision from internal appearance to explicit boundary modeling, the approach enables flexible instance definition—such as whole objects or subparts—via mask conditioning. A pixel-wise shallow MLP facilitates rapid training and fine-grained control. Experiments demonstrate superior few-shot generalization, robustness, and precise control over segmentation granularity on low-texture industrial components and food imagery.
This study addresses the challenge of effectively leveraging large-scale unlabeled images to improve object detection performance under limited annotation budgets. We systematically evaluate three prominent semi-supervised object detection methods—MixPL, Semi-DETR, and Consistent-Teacher—across MS-COCO, Pascal VOC, and a custom Beetle dataset, analyzing their trade-offs among accuracy, model size, and inference latency under varying labeling ratios. For the first time, we reveal consistent patterns of performance degradation as labeled data decreases, both on general-purpose and domain-specific datasets. Our empirical findings provide actionable insights and practical guidance for selecting appropriate semi-supervised approaches in resource-constrained scenarios.