object detection

Designs, implements, and evaluates models and systems that detect and localize instances of predefined object categories in images or video, producing outputs such as bounding boxes, class labels, and segmentation masks. Builds training and annotation pipelines, selects architectures and loss functions, and analyzes performance metrics (e.g., IoU, precision/recall), runtime, and robustness to occlusion, scale, and scene variation.

objectdetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Object Learning and Robust 3D Reconstruction

Apr 22, 2025
SS
Sara Sabour
🏛️ University of Toronto

This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.

3D dynamic object detection via geometric consistencyRobust 3D modeling with transient object masksUnsupervised 2D object segmentation using motion cues

To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.

Automate video annotation for computer vision modelsDevelop efficient video tracking and segmentation toolReduce time and resources for labeled data generation

Approximate Size Targets Are Sufficient for Accurate Semantic Segmentation

Mar 10, 2025
XF
Xingye Fan
🏛️ University of Waterloo

This work addresses weakly supervised semantic segmentation using only image-level class labels. We propose a novel paradigm that leverages approximate relative object size distributions—as opposed to pixel-level masks—as the weak supervision signal. Our method builds upon standard segmentation architectures (e.g., DeepLab, FCN) and introduces a zero-avoiding KL divergence loss to directly align predicted size distributions with coarse-grained size annotations, either human-provided or synthetically generated, without architectural modifications or multi-stage training. We provide the first theoretical analysis and empirical validation demonstrating that coarse-grained size distribution information alone is sufficient to achieve segmentation performance approaching fully supervised accuracy. On PASCAL VOC, our approach achieves state-of-the-art weakly supervised performance, with several classes even surpassing fully supervised baselines. Moreover, it exhibits strong generalization and robustness to annotation noise on COCO and medical imaging benchmarks.

Demonstrates accurate semantic segmentation using approximate size targets.Introduces zero-avoiding KL-divergence loss for comparable segmentation accuracy.Validates robustness of standard networks to size target errors.

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

Traditional object detection suffers from heavy reliance on labor-intensive manual annotations, poor generalization, and limited adaptability to novel categories and dynamic environments. To address these challenges, this work proposes an end-to-end fully automated detection pipeline. Methodologically, it introduces— for the first time—a unified framework integrating CLIP-driven open-vocabulary localization, diffusion model–enhanced feature representation, uncertainty-aware pseudo-label filtering, and an interactive human verification mechanism. This enables zero-shot category extension and closed-loop optimization with controllable annotation quality. Built upon a fine-tuned YOLOv8 backbone, the method achieves 92% of the full-supervision state-of-the-art mAP on COCO and LVIS using only 15% of the manual annotations required by conventional approaches. The proposed pipeline significantly reduces annotation cost while substantially improving cross-domain generalization capability.

Adapts to diverse environments and novel object categoriesAutomates end-to-end object detection without manual labelingImproves accuracy from data collection to model training

Latest Papers

What's happening recently
View more

This work proposes a training-free object detection method tailored for scenarios involving minor data variations where model training and annotation are impractical, such as GUI automation testing. By leveraging a segmentation foundation model—e.g., SAM—to generate image segments and integrating classical feature engineering for object classification, the approach rapidly adapts to new targets or interface changes without any training or labeled data. Evaluated on an in-vehicle navigation icon detection task, the method achieves performance comparable to learning-based detectors like YOLO, while entirely eliminating the need for model training. This significantly reduces deployment time and cost, demonstrating strong practical utility through its efficiency and adaptability.

foundation modelGUI testingicon detection

Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.

annotation structureobject detectionrepresentation robustness

This work addresses the challenge of few-shot object detection in industrial settings, where new objects are frequently introduced and labeled data are scarce. The authors propose a training-free detection method that leverages a vision foundation model to construct category prototypes from only a few reference images and integrates a segmentation model to generate region proposals. Detection is achieved through a decoupled prototype matching mechanism that enables efficient feature alignment. The approach allows rapid deployment for novel categories using just a handful of reference images, without requiring CAD models or large-scale annotations. Experiments on three industrial datasets demonstrate a 6.9% average precision (AP) improvement over the current best training-free methods, highlighting its strong practical potential.

few-shot object detectionindustrial object detectionlimited labeled data

This work proposes a novel turbo inference strategy that enables bidirectional iterative refinement between detection and segmentation during inference, without requiring retraining. Challenging the conventional unidirectional “detect-then-segment” paradigm in top-down instance segmentation, the method introduces a closed-loop interaction mechanism through turbo detection and segmentation heads, coupled with cross-task feature fusion. This design dynamically leverages complementary information from both tasks while preserving the top-down architecture. Extensive experiments demonstrate significant improvements in both detection and segmentation accuracy on COCO, iFLYTEK, and Cityscapes benchmarks, achieving a favorable trade-off between performance and inference efficiency.

detect-then-segment paradigminstance segmentationobject detection

This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.

cross-domain transferabilityfilter pipelinesmonitoring systems

Hot Scholars

JE

Jon E. Froehlich

Professor, Allen School of Computer Science, University of Washington
HCIHuman-Centered AIAccessibilityAugmented Reality
DT

Dzmitry Tsetserukou

Associate Professor, Skolkovo Institute of Science and Technology (Skoltech)
RoboticsHapticsUAV SwarmAI
RY

Rui-Yang Ju

National Taiwan University
Computer VisionHandwriting RecognitionDocument AnalysisMedical Image Processing
RG

Ross Greer

University of California Merced
Artificial IntelligenceMachine VisionAutonomous DrivingHuman-Robot Interaction
CL

Chu Li

PhD student, University of Washington
human-computer interactionurban planning