Score
Designs, implements, and evaluates models and systems that detect and localize instances of predefined object categories in images or video, producing outputs such as bounding boxes, class labels, and segmentation masks. Builds training and annotation pipelines, selects architectures and loss functions, and analyzes performance metrics (e.g., IoU, precision/recall), runtime, and robustness to occlusion, scale, and scene variation.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
This work addresses weakly supervised semantic segmentation using only image-level class labels. We propose a novel paradigm that leverages approximate relative object size distributions—as opposed to pixel-level masks—as the weak supervision signal. Our method builds upon standard segmentation architectures (e.g., DeepLab, FCN) and introduces a zero-avoiding KL divergence loss to directly align predicted size distributions with coarse-grained size annotations, either human-provided or synthetically generated, without architectural modifications or multi-stage training. We provide the first theoretical analysis and empirical validation demonstrating that coarse-grained size distribution information alone is sufficient to achieve segmentation performance approaching fully supervised accuracy. On PASCAL VOC, our approach achieves state-of-the-art weakly supervised performance, with several classes even surpassing fully supervised baselines. Moreover, it exhibits strong generalization and robustness to annotation noise on COCO and medical imaging benchmarks.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
Traditional object detection suffers from heavy reliance on labor-intensive manual annotations, poor generalization, and limited adaptability to novel categories and dynamic environments. To address these challenges, this work proposes an end-to-end fully automated detection pipeline. Methodologically, it introduces— for the first time—a unified framework integrating CLIP-driven open-vocabulary localization, diffusion model–enhanced feature representation, uncertainty-aware pseudo-label filtering, and an interactive human verification mechanism. This enables zero-shot category extension and closed-loop optimization with controllable annotation quality. Built upon a fine-tuned YOLOv8 backbone, the method achieves 92% of the full-supervision state-of-the-art mAP on COCO and LVIS using only 15% of the manual annotations required by conventional approaches. The proposed pipeline significantly reduces annotation cost while substantially improving cross-domain generalization capability.
This work proposes a training-free object detection method tailored for scenarios involving minor data variations where model training and annotation are impractical, such as GUI automation testing. By leveraging a segmentation foundation model—e.g., SAM—to generate image segments and integrating classical feature engineering for object classification, the approach rapidly adapts to new targets or interface changes without any training or labeled data. Evaluated on an in-vehicle navigation icon detection task, the method achieves performance comparable to learning-based detectors like YOLO, while entirely eliminating the need for model training. This significantly reduces deployment time and cost, demonstrating strong practical utility through its efficiency and adaptability.
Existing object detectors often learn task-driven features that rely on shortcut correlations, failing to adequately capture the underlying annotation structure, which limits their generalization, interpretability, and robustness under task shifts or sparse supervision. To address this, this work proposes an annotation-guided feature enhancement framework that explicitly integrates geometric annotation priors into feature learning for the first time. By constructing a dense spatial feature grid and injecting it into the backbone network—where it fuses with the feature pyramid—the method steers region proposal and detection heads toward representations better aligned with annotation structure. Evaluated on wildlife and remote sensing datasets, the approach significantly improves object focus, reduces background sensitivity, and demonstrates superior generalization and data efficiency in weakly supervised and unseen-task settings.
This work addresses the challenge of few-shot object detection in industrial settings, where new objects are frequently introduced and labeled data are scarce. The authors propose a training-free detection method that leverages a vision foundation model to construct category prototypes from only a few reference images and integrates a segmentation model to generate region proposals. Detection is achieved through a decoupled prototype matching mechanism that enables efficient feature alignment. The approach allows rapid deployment for novel categories using just a handful of reference images, without requiring CAD models or large-scale annotations. Experiments on three industrial datasets demonstrate a 6.9% average precision (AP) improvement over the current best training-free methods, highlighting its strong practical potential.
This work proposes a novel turbo inference strategy that enables bidirectional iterative refinement between detection and segmentation during inference, without requiring retraining. Challenging the conventional unidirectional “detect-then-segment” paradigm in top-down instance segmentation, the method introduces a closed-loop interaction mechanism through turbo detection and segmentation heads, coupled with cross-task feature fusion. This design dynamically leverages complementary information from both tasks while preserving the top-down architecture. Extensive experiments demonstrate significant improvements in both detection and segmentation accuracy on COCO, iFLYTEK, and Cityscapes benchmarks, achieving a favorable trade-off between performance and inference efficiency.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.