two-stage detection and segmentation

Designs, implements, and evaluates cascaded coarse‑to‑fine pipelines that first detect candidate foreground/object regions and then classify or segment those regions, including architectures that crop or mask regions to pass to a second-stage segmenter or classifier. Builds and analyzes background modeling, background removal/suppression, and foreground‑background separation components to reduce sensitivity to background noise, handle out‑of‑distribution backgrounds, and speed processing of large datasets by limiting computation to detected regions.

two-stagedetectionandsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Automated Video Segmentation Machine Learning Pipeline

Jul 09, 2025
JM
Johannes Merz
🏛️ Image Engine Design Inc.

To address the inefficiency, time consumption, and high resource demands of manual mask generation in visual effects (VFX) production, this paper proposes a text-prompt-driven automated video segmentation pipeline. The method integrates text-guided instance segmentation for flexible object localization, fine-grained frame-wise semantic segmentation, and temporally consistent video object tracking, deployed via lightweight containerization to ensure seamless integration with artist workflows. Key contributions include: (i) the first end-to-end incorporation of text prompting into VFX-grade video segmentation; (ii) a multi-stage collaborative architecture ensuring inter-frame segmentation stability; and (iii) structured, editable outputs—including mask sequences, confidence maps, and trajectory metadata. Experiments demonstrate that the system reduces pre-compositing preparation time by 72% on average and decreases manual intervention by 89%, significantly enhancing automation and productivity across VFX pipelines.

Automates slow mask generation in VFX productionEnsures temporally consistent video segmentation masksReduces manual effort for preliminary composites

Small-object detection performance is hindered by fragmented optimization across stages in conventional pipeline-based detectors. To address this, we propose PLUSNet, an end-to-end co-optimization framework introducing the novel “Purify–Label–Utilize” paradigm. Specifically, we design a hierarchical feature purifier to suppress noise; develop a multi-criterion dynamic label assignment mechanism to improve positive/negative sample quality; and introduce a frequency-domain decoupled detection head for fine-grained feature modeling. All modules are lightweight, modular, and seamlessly integrate with mainstream detectors. Extensive experiments on MS COCO, VisDrone, and other benchmarks demonstrate consistent and significant gains in small-object AP (+3.2–5.8 points), validating the effectiveness of joint upstream-downstream optimization and strong generalizability across diverse scenarios and architectures.

Enhancing downstream task performance effectivelyImproving feature purification and sample labelingOptimizing small object detection pipeline holistically

This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.

cross-domain transferabilityfilter pipelinesmonitoring systems

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

On Efficient Variants of Segment Anything Model: A Survey

Oct 07, 2024
XS
Xiaorui Sun
🏛️ UESTC | Lancaster University | Tongji University

While the Segment Anything Model (SAM) exhibits strong generalization capability, its substantial computational overhead hinders deployment on resource-constrained edge devices. This work presents a systematic survey of efficient SAM variants tailored for edge deployment. We introduce the first unified evaluation framework spanning diverse hardware platforms—including CPU, GPU, and Edge TPU—and conduct joint accuracy–latency–memory benchmarking on COCO and SA-1B. Our analysis categorizes acceleration techniques along six technical axes: model pruning, knowledge distillation, lightweight attention mechanisms, quantization, module substitution, and hardware-aware compilation—characterizing their Pareto-optimal trade-offs. The core contributions are: (1) an open-source, fully reproducible edge-SAM benchmark; and (2) empirical insights into the applicability domains and fundamental accuracy-efficiency trade-offs of each acceleration strategy—providing both theoretical foundations and practical guidelines for designing lightweight vision foundation models.

Addressing high computational demands of Segment Anything ModelEnhancing SAM efficiency for resource-limited environmentsSurveying acceleration techniques for SAM variants

Latest Papers

What's happening recently
View more

Existing CLIP-based foreground-background decomposition methods for few-shot out-of-distribution (OOD) detection employ a uniform background suppression strategy and overlook semantically ambiguous local regions within the foreground, thereby limiting performance. To address these limitations, this work proposes a plug-and-play framework that refines semantic control over image regions through three key components: foreground-background decomposition, entropy-weighted adaptive background suppression, and confusion-aware foreground correction based on local semantic similarity analysis. By moving beyond the conventional practice of treating all background regions equally and ignoring intra-foreground semantic ambiguities, the proposed method achieves significant performance gains over existing foreground-background approaches across multiple benchmarks, demonstrating its effectiveness and generalizability.

background suppressionCLIPconfusable foreground

Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.

binary segmentationevaluation metricsmetric decomposition

This study addresses the challenges of detail loss and dataset discrepancies in satellite image time series crop segmentation by proposing PAtteRNS, a hybrid Transformer-convolutional model. Its core innovation lies in the first parallel dimension self-attention architecture tailored for Sentinel-2 multispectral data, which independently decouples attention computation across temporal, spectral, and spatial dimensions, achieving fully factorized attention while significantly reducing computational complexity. Experimental results demonstrate that PAtteRNS surpasses existing state-of-the-art methods on the PASTIS and MTLCC datasets, exhibiting particularly notable advantages in parcel boundary delineation quality. Furthermore, this work reveals the potential impact of dataset grouping deficiencies on model performance.

Cropland SegmentationDataset DisparityParcel Delineation

Existing vision Mixture-of-Experts (MoE) approaches perform routing at the image or patch level, which struggles to align with the instance-centric nature of object detection. This work proposes a hierarchical instance-conditioned MoE architecture that introduces a two-stage routing mechanism—operating at both scene and instance levels—within a DETR-style detector, achieving fine-grained expert assignment aligned with instance queries for the first time. The method employs a lightweight scene router and an instance router that jointly balance sparse computation with the heterogeneity of individual instances. Experiments demonstrate that the model outperforms dense DINO baselines and simplified routing variants on COCO, significantly enhancing small object detection performance and offering preliminary evidence of functional specialization among experts.

expert specializationinstance-level routingMixture-of-Experts

Hot Scholars

JP

Jigen Peng

Guangzhou University
Sparse Infroamtion ProcessingNonlinear Functional AnalysisMotion Detection
QW

Qi Wang

Northwestern Polytechnical University
Computer visionPattern recognitionMachine learningRemote sensing
MI

Mariko Isogawa

Keio University
computer visionmachine learningaugmented realityimage processing
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing