Score
Designs, implements, and evaluates models, algorithms, and end-to-end pipelines that process visual data (images and video) to detect and localize objects, perform semantic and instance segmentation, and track multiple objects over time. Builds and analyzes methods for image processing, depth and pose estimation, visual grounding, segmentation modeling/analysis, scene understanding, and related computer vision techniques and algorithms.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
This work addresses the challenge of achieving real-time performance, high accuracy, and energy efficiency in embedded vision systems operating on resource-constrained hardware. The authors propose an algorithm-hardware co-design methodology tailored for DSP/FPGA platforms, optimizing edge, corner, and blob detection operators through hardware-aware algorithmic refinements and quantization techniques. To further enhance throughput without compromising image quality, the approach incorporates inter-frame redundancy elimination and adaptive frame averaging strategies. Experimental results demonstrate that, compared to conventional solutions, the proposed method delivers significantly improved processing speed and energy efficiency, enabling scalable and highly effective real-time embedded vision across diverse applications such as automotive systems, surveillance, and robotics.
To address the challenges of heavy reliance on dense pixel-level annotations and poor adaptability to diverse industrial conditions in scrap sorting—particularly for foreign object detection and segmentation—this paper proposes a novel weakly supervised learning paradigm termed “before-after supervision,” leveraging image differences before and after human intervention. Methodologically, we design an end-to-end semantic segmentation framework integrating differential feature modeling and multi-view consistency constraints, enabling unified evaluation of diverse weak supervision strategies. Key contributions include: (1) the first formulation of operator removal actions as natural weak supervision signals; (2) the release of WS², the first multi-view, high-resolution weakly supervised dataset specifically designed for industrial scrap sorting; and (3) empirical validation on WS² demonstrating that multiple weakly supervised methods achieve over 90% of fully supervised performance, substantially reducing annotation costs while maintaining practical deployability.
To address insufficient geometric-semantic joint understanding of robots in unstructured environments, this paper proposes an end-to-end modular RGB-D understanding pipeline. Our method introduces a novel hybrid mask generation mechanism integrating SAM2 with a lightweight semantic classifier for pixel-level semantic segmentation and instance awareness; incorporates ReID-enhanced cross-frame human tracking and semantic-weighted TSDF point cloud fusion to ensure geometric fidelity and semantic consistency; and outputs structured scene representations in USD format. Evaluated on ADE20K, our approach achieves 47.0% mIoU—surpassing SegFormer and OneFormer in boundary accuracy—and attains a reconstruction error of only 25.3 mm. It runs 1.81× faster than prior methods and has been validated for deployment feasibility on real-world Kinect data.
This work addresses motion-induced artifacts in unsupervised 3D reconstruction of dynamic scenes caused by moving foreground objects. To tackle this, we propose a novel method that jointly models unsupervised 2D object decomposition and 3D geometric consistency. Our approach introduces FlowCapsules—a flow-guided network for unsupervised foreground segmentation—and a transient-object-mask-driven robust optimization kernel that detects and excludes dynamic objects under multi-view geometric consistency constraints. This kernel subsequently guides weighted bundle adjustment and NeRF training. Crucially, our method is the first to jointly learn object-centric representations and scene-level geometry without any annotations or controlled capture conditions. Experiments demonstrate significant improvements in both SfM and NeRF reconstruction accuracy on casually captured dynamic scenes, effectively eliminating motion artifacts. The framework advances explicit object-aware 3D understanding in open-world vision applications.
This work proposes a training-free object detection method tailored for scenarios involving minor data variations where model training and annotation are impractical, such as GUI automation testing. By leveraging a segmentation foundation model—e.g., SAM—to generate image segments and integrating classical feature engineering for object classification, the approach rapidly adapts to new targets or interface changes without any training or labeled data. Evaluated on an in-vehicle navigation icon detection task, the method achieves performance comparable to learning-based detectors like YOLO, while entirely eliminating the need for model training. This significantly reduces deployment time and cost, demonstrating strong practical utility through its efficiency and adaptability.
This work proposes SuperCam, a novel camera architecture designed to address the inefficiency of conventional cameras in resource-constrained settings, where they generate excessive redundant data that hinders downstream visual tasks. SuperCam uniquely integrates adaptive superpixel segmentation directly into the hardware pipeline, enabling online, lightweight data compression and preservation of critical visual information at the point of capture. By synergizing this design with edge computing, the system substantially reduces memory footprint and bandwidth requirements. Experimental results demonstrate that SuperCam consistently outperforms traditional approaches across multiple vision tasks—including semantic segmentation, object detection, and monocular depth estimation—thereby validating its feasibility and superiority for efficient perception under stringent resource limitations.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.
This work addresses the high cost of deploying traditional warehouse vision systems, which typically require repeated data collection, annotation, and retraining for each new environment. To overcome this limitation, the authors propose a generalizable deployment framework that eliminates the need for on-site retraining. By optimizing camera placement, implementing strategic image triggering mechanisms, selecting robust foundation models, and employing effective ensemble strategies, the system is trained once in the lab using only laboratory-collected data. The study demonstrates, for the first time, that such a model can successfully generalize across diverse real-world warehouse settings. Specifically, in the task of detecting fork anomalies in vertical material handling systems, deployment is reduced to simply installing cameras, capturing images, and directly applying the pre-trained model—bypassing costly annotation and retraining cycles and significantly lowering real-world deployment costs.
This study addresses the challenge of accurately segmenting irregular and densely packed components in electronic waste by presenting the first systematic comparison between the general-purpose foundation model SAM2 and the lightweight task-specific model YOLOv8. Leveraging a newly curated dataset of 1,456 high-quality annotated images and diverse data augmentation strategies, experiments demonstrate that YOLOv8 significantly outperforms SAM2, achieving an mAP50 of 98.8% and an mAP50–95 of 85%, with notably more precise boundary delineation. Although SAM2 offers greater structural flexibility, it suffers from mask overlap and contour inconsistency issues. The findings underscore that off-the-shelf vision foundation models require targeted adaptation to suit industrial recycling applications. The authors publicly release the dataset and benchmark framework to support further research in this domain.